Skip to content
ModelCap

Agent ranking

Best AI models for agents

Language models measured on at least two public agent boards (LMArena Agent, BFCL, τ²-bench), ordered by the agent component of the ModelCap Index. Models with a single agent result follow in their own tier, ordered by that board, so one result never outranks two.

Snapshot as of 9 September 2026

GPT-5.6 Sol from OpenAI leads the best-for-agents list as of 9 September 2026 with LMArena Agent 7.7% · AA τ²-Bench Telecom 85.1% across 2 public agent boards, holding ModelCap Index position #5; 29 models qualify.

Models on this list
29
With openly downloadable weights
9
With a listed API price
29

Best AI models for agents (4)

Every row is a current, canonical language model with a public position on the ModelCap Index; the list filters by a published fact and orders by a published figure.

Best AI models for agents: current, canonical language models with a measured public ModelCap Index position and admitted results on at least two public agent-family boards, ordered by the agent component of the ModelCap Index (ties broken by Index rank)
#ModelAgent boardsModelCap IndexInput $/1MOutput $/1MContextWeights
1GPT-5.6 SolOpenAILMArena Agent 7.7% · AA τ²-Bench Telecom 85.1%#583.0$2.00$10.001MAPI only
2GPT-5.5OpenAILMArena Agent 5.3% · AA τ²-Bench Telecom 93.9%#979.8$5.00$30.001MAPI only
3Kimi K3Moonshot AILMArena Agent 6.6%#682.6$3.00$15.001MRestricted license
4MiMo-V2.5-ProXiaomiLMArena Agent -5.1% · AA τ²-Bench Telecom 94.2%#2172.8$0.435$0.871MOpen weights

Measured on one agent board (25)

Ranked models with a single admitted agent result, ordered by that board's normalised score. One board is a narrower claim than two, so these rows sit apart from the list above and never mix into it.

Measured on one agent board: Ranked models with a single admitted agent result, ordered by that board's normalised score. One board is a narrower claim than two, so these rows sit apart from the list above and never mix into it.
#ModelAgent boardsModelCap IndexInput $/1MOutput $/1MContextWeights
1Claude Fable 5.1AnthropicLMArena Agent 14.5%#387.7$10.00$50.001MAPI only
2GPT-6 AstraOpenAILMArena Agent 12.5%#193.7$10.00$50.001MAPI only
3Claude Opus 5AnthropicLMArena Agent 11.4%#486.2$5.00$25.001MAPI only
4Claude Sonnet 5AnthropicLMArena Agent 6.3%#1875.6$2.00$10.001MAPI only
5Hy4 previewTencentLMArena Agent 5.4%#2670.4$0.834$2.501MOpen weights
6Gemini 3.8 FlashGoogleLMArena Agent 4.0%#782.2$0.75$3.751MAPI only
7Grok 4.6xAILMArena Agent 3.2%#1675.8$2.00$6.00500KAPI only
8GLM 5.3Z.aiLMArena Agent 2.6%#882.1$1.40$4.401.3MRestricted license
9GLM 5.3 FlashZ.aiLMArena Agent 2.0%#1178.7$0.075$0.251.3MOpen weights
10Qwen3.8 27BQwen#3565.2$0.42$3.001MOpen weights
11GPT-5.6 TerraOpenAILMArena Agent 1.5% · AA τ²-Bench Telecom 86.3%#1576.1$2.00$12.001MAPI only
12GPT-5.6 LunaOpenAILMArena Agent 1.1%#2570.5$0.20$1.201MAPI only
13Muse Glimmer 30BMeta#2868.4$0.30$1.10131KOpen weights
14Kimi K2.7 CodeMoonshot AIAA τ²-Bench Telecom 90.1%#4759.5$0.71$3.50262KRestricted license
15Gemma 4 31BGoogleAA τ²-Bench Telecom 65.5%#2769.7$0.09$0.34262KOpen weights
16Qwen3.7 PlusQwenLMArena Agent -5.2% · AA τ²-Bench Telecom 93.0%#2968.3$0.32$1.281MAPI only
17MiniMax M3MiniMaxLMArena Agent -5.3% · AA τ²-Bench Telecom 88.9%#3465.6$0.30$1.201MRestricted license
18Gemini 3.1 Pro PreviewGoogleLMArena Agent -5.3% · AA τ²-Bench Telecom 95.6%#1278.4$2.00$12.001MAPI only
19DeepSeek V3.2DeepSeekAA τ²-Bench Telecom 78.9%#5158.2$0.269$0.40164KOpen weights
20Inkling SmallThinking MachinesLMArena Agent -6.9%#6453.3$0.45$1.201MOpen weights
21GLM 5 TurboZ.aiAA τ²-Bench Telecom 98.5%#7350.4$1.20$4.00203KAPI only
22InklingThinking MachinesLMArena Agent -8.9%#4162.9$1.00$4.051MOpen weights
23Mistral Medium 3.5Mistral AILMArena Agent -9.3% · AA τ²-Bench Telecom 94.2%#5656.0$1.50$7.50262KAPI only
24Solar Pro 4UpstageLMArena Agent -11.3%#9342.3$0.03$0.12524KAPI only
25Gemini 3.5 Flash LiteGoogleLMArena Agent -13.4%#3067.1$0.30$2.501MAPI only

Best AI models for agents: common questions

Which model tops the best-for-agents list right now?

GPT-5.6 Sol (OpenAI), as of 9 September 2026, with LMArena Agent 7.7% · AA τ²-Bench Telecom 85.1% across 2 public agent boards and ModelCap Index position #5.

How is the best-for-agents list built?

It contains current, canonical language models with a measured public ModelCap Index position and admitted results on at least two public agent-family boards, ordered by the agent component of the ModelCap Index (ties broken by Index rank). A second tier, "Measured on one agent board", follows: Ranked models with a single admitted agent result, ordered by that board's normalised score. One board is a narrower claim than two, so these rows sit apart from the list above and never mix into it. Nothing is estimated for the list itself: the filter is a published fact and the order is a published figure, and every row links to the model page that shows its evidence.

How many models qualify for the best-for-agents list?

29 models qualify (4 in the primary tier, 25 measured on one board). The population is the live ModelCap Index, so a model appears here the moment it holds a public position and meets the filter.

Which model on this list ranks highest on the ModelCap Index?

GPT-6 Astra holds ModelCap Index position #1, the highest of any model in the list, and sits at #2 here by agent boards.

What is the cheapest model on the best-for-agents list?

Solar Pro 4 lists the lowest blended API price at $0.052 per 1M tokens ($0.03 input / $0.12 output).

Which is the best open-weight model on this list?

MiMo-V2.5-Pro from Xiaomi is the highest-placed model with openly downloadable weights (#4 here, ModelCap Index #21).

How current is the best-for-agents list?

It is rendered from the sealed dataset published 9 September 2026 and re-renders within a minute of each refresh; positions, prices and context figures are the ones shown on the live rankings at the same instant.

Explore further