Terminal-Bench 2.1 asks a model to complete real engineering work inside a Linux terminal. Artificial Analysis runs every model through the same fixed harness, so the pass rate compares models rather than the vendor agents that dominate the official board. This page lists every model ModelCap tracks with a published AA Terminal-Bench 2.1 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 2 September 2026
Claude Fable 5.1 from Anthropic leads the Terminal-Bench 2.1 (Artificial Analysis run) among the 83 tracked models with 91.4% (terminal-bench-2.1 configuration), per the source snapshot published 2 September 2026.
AA Terminal-Bench 2.1 ranks models by Terminal-Bench 2.1 task pass rate on one fixed harness. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as coding evidence alongside the other public boards; the methodology documents the weighting and the identity rules.
AA Terminal-Bench 2.1 leaderboard: common questions
What is the AA Terminal-Bench 2.1 benchmark?
Terminal-Bench 2.1 asks a model to complete real engineering work inside a Linux terminal. Artificial Analysis runs every model through the same fixed harness, so the pass rate compares models rather than the vendor agents that dominate the official board. It is published by Artificial Analysis.
Which AI model leads AA Terminal-Bench 2.1 right now?
Claude Fable 5.1 (Anthropic) holds the top AA Terminal-Bench 2.1 score among the models ModelCap tracks, at 91.4% as of the source snapshot published 2 September 2026.
How many models are ranked on the AA Terminal-Bench 2.1 leaderboard here?
83 tracked models have a published AA Terminal-Bench 2.1 result on ModelCap; the source board itself lists 226 entries. Each row shows the model's best evaluated configuration.
What is the best open-weight model on AA Terminal-Bench 2.1?
Qwen3.8 27B from Qwen is the highest-scoring model with openly downloadable weights on this board, at 79.8%.
Which model offers the best value on AA Terminal-Bench 2.1?
Among the ten highest-scoring priced models, Qwen3.8 Flash has the lowest listed output price at $0.47 per 1M tokens while scoring 86.1%.
How is AA Terminal-Bench 2.1 scored?
The board ranks models by Terminal-Bench 2.1 task pass rate on one fixed harness; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does AA Terminal-Bench 2.1 decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; AA Terminal-Bench 2.1 contributes as coding evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the AA Terminal-Bench 2.1 results?
The newest AA Terminal-Bench 2.1 publication ModelCap holds is dated 2 September 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.