Skip to content
ModelCap

Benchmark leaderboard

Terminal-Bench 2.1 (Artificial Analysis run)

Terminal-Bench 2.1 asks a model to complete real engineering work inside a Linux terminal. Artificial Analysis runs every model through the same fixed harness, so the pass rate compares models rather than the vendor agents that dominate the official board. This page lists every model ModelCap tracks with a published AA Terminal-Bench 2.1 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 2 September 2026

Claude Fable 5.1 from Anthropic leads the Terminal-Bench 2.1 (Artificial Analysis run) among the 83 tracked models with 91.4% (terminal-bench-2.1 configuration), per the source snapshot published 2 September 2026.

Tracked models with a result
83
Entries on the source board
226
Open-weight models listed
25

AA Terminal-Bench 2.1 standings (83 models)

Sorted by published score, one row per model. Source: Artificial Analysis.

Terminal-Bench 2.1 (Artificial Analysis run): every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Claude Fable 5.1Anthropic91.4%terminal-bench-2.1#1 / 226#2$50.001MAPI only2 September 2026
2Claude Opus 5Anthropic89.1%terminal-bench-2.1#5 / 226#3$25.001MAPI only2 September 2026
3Grok 4.6xAI88.4%terminal-bench-2.1#6 / 226#16$6.00500KAPI only2 September 2026
4GPT-5.6 SolOpenAI88.0%terminal-bench-2.1#7 / 226#5$10.001MAPI only2 September 2026
5GPT-5.6 TerraOpenAI88.0%terminal-bench-2.1#7 / 226#13$12.001MAPI only2 September 2026
6Gemini 3.8 FlashGoogle87.6%terminal-bench-2.1#12 / 226#11$3.751MAPI only2 September 2026
7Qwen3.8 FlashQwen86.1%terminal-bench-2.1#15 / 226#17$0.471MRestricted license2 September 2026
8Gemini 3.7 FlashGoogle85.8%terminal-bench-2.1#18 / 226$3.751MAPI only2 September 2026
9Kimi K3Moonshot AI85.0%terminal-bench-2.1#19 / 226#6$15.001MRestricted license2 September 2026
10Claude Fable 5Anthropic84.6%terminal-bench-2.1#21 / 226#1$50.001MAPI only2 September 2026
11Claude Opus 4.8Anthropic84.6%terminal-bench-2.1#21 / 226$25.001MAPI only2 September 2026
12GPT-5.5OpenAI84.3%terminal-bench-2.1#23 / 226#9$30.001MAPI only2 September 2026
13Claude Opus 4.7Anthropic83.1%terminal-bench-2.1#28 / 226$25.001MAPI only2 September 2026
14Grok 4.5xAI81.6%terminal-bench-2.1#32 / 226$6.00500KAPI only2 September 2026
15Qwen3.8 MaxQwen81.3%terminal-bench-2.1#33 / 226#7$6.001MAPI only2 September 2026
16GPT-5.6 LunaOpenAI80.9%terminal-bench-2.1#34 / 226#18$1.201MAPI only2 September 2026
17Claude Sonnet 5Anthropic80.5%terminal-bench-2.1#35 / 226#15$10.001MAPI only2 September 2026
18Muse Spark 1.2Meta80.1%terminal-bench-2.1#37 / 226$4.251MAPI only2 September 2026
19Qwen3.8 27BQwen79.8%terminal-bench-2.1#39 / 226#26$2.551MOpen weights2 September 2026
20Gemini 3.5 FlashGoogle78.7%terminal-bench-2.1#42 / 226$9.001MAPI only2 September 2026
21GPT-5.4OpenAI78.3%terminal-bench-2.1#45 / 226$15.001MAPI only2 September 2026
22GLM 5.2Z.ai77.9%terminal-bench-2.1#47 / 226$3.041MOpen weights2 September 2026
23Muse Spark 1.1Meta77.9%terminal-bench-2.1#47 / 226$4.251MAPI only2 September 2026
24Gemini 3.6 FlashGoogle77.5%terminal-bench-2.1#50 / 226$3.751MAPI only2 September 2026
25Qwen3.7 MaxQwen74.5%terminal-bench-2.1#57 / 226$4.421MAPI only2 September 2026
26Gemini 3.1 Pro PreviewGoogle73.8%terminal-bench-2.1#60 / 226#8$12.001MAPI only2 September 2026
27KAT-Coder-Pro V2Kuaishou70.0%terminal-bench-2.1#64 / 226#57$1.20262KAPI only2 September 2026
28Nex-N2-ProNex AGI67.8%terminal-bench-2.1#69 / 226#33$1.00262KOpen weights2 September 2026
29Kimi K2.7 CodeMoonshot AI67.4%terminal-bench-2.1#70 / 226#30$3.40262KRestricted license2 September 2026
30Kimi K2.6Moonshot AI65.9%terminal-bench-2.1#73 / 226$4.00262KRestricted license2 September 2026
31MiniMax M3MiniMax65.2%terminal-bench-2.1#75 / 226#25$1.201MRestricted license2 September 2026
32Hy3Tencent64.4%terminal-bench-2.1#79 / 226#20$0.528262KOpen weights2 September 2026
33GLM 5.1Z.ai61.8%terminal-bench-2.1#84 / 226$3.04205KOpen weights2 September 2026
34Qwen3.6 PlusQwen61.4%terminal-bench-2.1#87 / 226$1.951MAPI only2 September 2026
35Qwen3.7 PlusQwen61.0%terminal-bench-2.1#88 / 226#21$1.281MAPI only2 September 2026
36GPT-5.4 NanoOpenAI60.7%terminal-bench-2.1#90 / 226#47$1.25400KAPI only2 September 2026
37Qwen3.6 27BQwen60.7%terminal-bench-2.1#90 / 226$3.60262KOpen weights2 September 2026
38GPT-5.4 MiniOpenAI59.2%terminal-bench-2.1#93 / 226#22$4.50400KAPI only2 September 2026
39Solar Pro 4Upstage57.3%terminal-bench-2.1#94 / 226#58$0.12524KAPI only2 September 2026
40Ling-3.0-flashInclusionAI55.4%terminal-bench-2.1#98 / 226#49$0.063262KOpen weights2 September 2026
41InklingThinking Machines55.1%terminal-bench-2.1#100 / 226#27$4.051MOpen weights2 September 2026
42Inkling SmallThinking Machines55.1%terminal-bench-2.1#100 / 226#42$1.201MOpen weights2 September 2026
43Nemotron 3 UltraNVIDIA53.9%terminal-bench-2.1#102 / 226#48$2.40262KRestricted license2 September 2026
44Gemini 3.5 Flash LiteGoogle53.6%terminal-bench-2.1#103 / 226#23$2.501MAPI only2 September 2026
45GPT-5.1OpenAI52.4%terminal-bench-2.1#105 / 226$10.00400KAPI only2 September 2026
46Qwen3.5 397B A17BQwen51.3%terminal-bench-2.1#109 / 226#28$3.50262KOpen weights2 September 2026
47Mistral Medium 3.5Mistral AI50.6%terminal-bench-2.1#111 / 226#41$7.50262KAPI only2 September 2026
48LongCat 2.0Meituan50.2%terminal-bench-2.1#112 / 226#54$1.201MOpen weights2 September 2026
49Qwen3.5-122B-A10BQwen47.6%terminal-bench-2.1#115 / 226#43$2.40262KOpen weights2 September 2026
50Kimi K2.5Moonshot AI45.7%terminal-bench-2.1#118 / 226$2.25262KRestricted license2 September 2026
51GLM 4.7Z.ai45.3%terminal-bench-2.1#119 / 226$1.75205KOpen weights2 September 2026
52Qwen3.6 35B A3BQwen44.9%terminal-bench-2.1#120 / 226#66$0.90262KOpen weights2 September 2026
53Gemma 4 31BGoogle43.4%terminal-bench-2.1#124 / 226#31$0.34262KOpen weights2 September 2026
54Grok 4.3xAI39.7%terminal-bench-2.1#130 / 226$2.501MAPI only2 September 2026
55Step 3.7 FlashStepFun39.3%terminal-bench-2.1#131 / 226#75$1.15262KOpen weights2 September 2026
56Gemma 4 26B A4BGoogle39.0%terminal-bench-2.1#132 / 226#38$0.34262KOpen weights2 September 2026
57Nemotron 3 SuperNVIDIA38.6%terminal-bench-2.1#135 / 226#79$0.401MRestricted license2 September 2026
58Qwen3 Coder NextQwen38.2%terminal-bench-2.1#136 / 226#91$0.80262KOpen weights2 September 2026
59North Mini CodeCohere35.6%terminal-bench-2.1#138 / 226#96Free256KOpen weights2 September 2026
60GPT-5OpenAI35.2%terminal-bench-2.1#139 / 226$10.00400KAPI only2 September 2026
61Gemini 3.1 Flash Lite PreviewGoogle31.1%terminal-bench-2.1#143 / 226$1.501MAPI only2 September 2026
62Qwen3.5-9BQwen29.2%terminal-bench-2.1#148 / 226#87$0.15262KOpen weights2 September 2026
63Gemini 2.5 ProGoogle28.5%terminal-bench-2.1#150 / 226$10.001MAPI only2 September 2026
64Ling 3.0 tinyInclusionAI27.7%terminal-bench-2.1#151 / 226#83Open weights2 September 2026
65Mercury 2Inception Labs27.3%terminal-bench-2.1#152 / 226#85$0.75128KAPI only2 September 2026
66granite 4.2 30bIBM Granite26.6%terminal-bench-2.1#154 / 226#86131KOpen weights2 September 2026
67Nemotron 3.5 LightningNVIDIA24.3%terminal-bench-2.1#158 / 226#84$0.20262KRestricted license2 September 2026
68Trinity Large ThinkingArcee AI20.6%terminal-bench-2.1#165 / 226#78$0.80262KRestricted license2 September 2026
69Granite 4.2 8BIBM Granite18.4%terminal-bench-2.1#169 / 226#176$0.15131KOpen weights2 September 2026
70granite 4.2 3bIBM Granite13.9%terminal-bench-2.1#175 / 226#178131KOpen weights2 September 2026
71Mistral Medium 3.1Mistral AI13.9%terminal-bench-2.1#175 / 226$2.00131KAPI only2 September 2026
72Solar Pro 3Upstage12.0%terminal-bench-2.1#182 / 226$0.60131KAPI only2 September 2026
73GPT-4.1 MiniOpenAI10.1%terminal-bench-2.1#186 / 226$1.601MAPI only2 September 2026
74Llama 4 MaverickMeta7.9%terminal-bench-2.1#189 / 226#102$0.6961MGated access2 September 2026
75GPT-4o-miniOpenAI5.6%terminal-bench-2.1#194 / 226#211$0.60128KAPI only2 September 2026
76Gemma 3 27BGoogle4.5%terminal-bench-2.1#199 / 226#90$0.45131KGated access2 September 2026
77LFM2.5-2.6BLiquid4.5%terminal-bench-2.1#199 / 226#183Free66KRestricted license2 September 2026
78GPT-4.1 NanoOpenAI3.7%terminal-bench-2.1#204 / 226$0.401MAPI only2 September 2026
79GPT-5 MiniOpenAI3.7%terminal-bench-2.1#204 / 226$2.00400KAPI only2 September 2026
80Llama 4 ScoutMeta3.7%terminal-bench-2.1#204 / 226#172$0.301.3MGated access2 September 2026
81Granite 4.1 8BIBM Granite3.4%terminal-bench-2.1#208 / 226$0.10131KOpen weights2 September 2026
82Gemma 3 4BGoogle0.4%terminal-bench-2.1#218 / 226#195$0.10131KGated access2 September 2026
83Gemma 3 12BGoogle0.0%terminal-bench-2.1#222 / 226#175$0.15131KGated access2 September 2026

How ModelCap uses AA Terminal-Bench 2.1

AA Terminal-Bench 2.1 ranks models by Terminal-Bench 2.1 task pass rate on one fixed harness. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as coding evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

AA Terminal-Bench 2.1 leaderboard: common questions

What is the AA Terminal-Bench 2.1 benchmark?

Terminal-Bench 2.1 asks a model to complete real engineering work inside a Linux terminal. Artificial Analysis runs every model through the same fixed harness, so the pass rate compares models rather than the vendor agents that dominate the official board. It is published by Artificial Analysis.

Which AI model leads AA Terminal-Bench 2.1 right now?

Claude Fable 5.1 (Anthropic) holds the top AA Terminal-Bench 2.1 score among the models ModelCap tracks, at 91.4% as of the source snapshot published 2 September 2026.

How many models are ranked on the AA Terminal-Bench 2.1 leaderboard here?

83 tracked models have a published AA Terminal-Bench 2.1 result on ModelCap; the source board itself lists 226 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on AA Terminal-Bench 2.1?

Qwen3.8 27B from Qwen is the highest-scoring model with openly downloadable weights on this board, at 79.8%.

Which model offers the best value on AA Terminal-Bench 2.1?

Among the ten highest-scoring priced models, Qwen3.8 Flash has the lowest listed output price at $0.47 per 1M tokens while scoring 86.1%.

How is AA Terminal-Bench 2.1 scored?

The board ranks models by Terminal-Bench 2.1 task pass rate on one fixed harness; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does AA Terminal-Bench 2.1 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; AA Terminal-Bench 2.1 contributes as coding evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the AA Terminal-Bench 2.1 results?

The newest AA Terminal-Bench 2.1 publication ModelCap holds is dated 2 September 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Explore further