Skip to content
ModelCap

Benchmark leaderboard

LMArena Agent leaderboard

LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. This page lists every model ModelCap tracks with a published LMArena Agent result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 30 September 2026

Claude Fable 5.1 from Anthropic leads the LMArena Agent leaderboard among the 40 tracked models with 14.6% (Max configuration), per the source snapshot published 30 September 2026.

Tracked models with a result
40
Entries on the source board
46
Open-weight models listed
8

LMArena Agent standings (40 models)

Sorted by published score, one row per model. Source: LMArena.

LMArena Agent leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Claude Fable 5.1Anthropic14.6%12.6%–16.5%Max#1 / 46#4$50.001MAPI only30 September 2026
2Claude Opus 5.5Anthropic13.8%11.5%–16.0%High#2 / 46#2$20.001MAPI only30 September 2026
3GPT-6 AstraOpenAI12.2%9.9%–14.5%Max#3 / 46#8$50.001MAPI only30 September 2026
4GPT-6 SolOpenAI10.6%8.1%–13.2%Max#4 / 46—$10.001MAPI only30 September 2026
5Claude Opus 5Anthropic8.8%7.3%–10.3%High+1 more evaluated#5 / 46—$25.001MAPI only30 September 2026
6Claude Fable 5Anthropic8.4%7.1%–9.7%High#7 / 46—$50.001MAPI only30 September 2026
7GPT-5.6 SolOpenAI7.1%5.7%–8.4%X-High#9 / 46—$10.001MAPI only30 September 2026
8Claude Opus 4.8Anthropic6.9%5.5%–8.3%High#10 / 46—$25.001MAPI only30 September 2026
9Claude Sonnet 5Anthropic4.8%3.1%–6.5%High#11 / 46—$10.001MAPI only30 September 2026
10Hy4 previewTencent4.2%3.4%–5.1%Default#12 / 46#26$2.251MOpen weights30 September 2026
11Muse Spark 1.3Meta4.2%3.6%–4.8%Max#13 / 46#5$4.251MAPI only30 September 2026
12GPT-5.5OpenAI4.1%3.1%–5.2%X-High+1 more evaluated#14 / 46#14$30.001MAPI only30 September 2026
13DeepSeek V4.1 FlashDeepSeek4.0%3.5%–4.5%Max#15 / 46#15$0.501MOpen weights30 September 2026
14Kimi K3Moonshot AI4.0%3.5%–4.5%Max#16 / 46#10$10.001MRestricted license30 September 2026
15Grok 4.7xAI3.6%2.1%–5.2%X-High#17 / 46#29$6.00500KAPI only30 September 2026
16GLM 5.2Z.ai3.6%2.9%–4.4%Max#18 / 46—$3.991MOpen weights30 September 2026
17Gemini 3.8 FlashGoogle3.0%2.1%–3.8%High#19 / 46#9$3.751MAPI only30 September 2026
18Qwen3.8 Max (0803)Qwen2.8%2.2%–3.5%Default#20 / 46—$6.001MRestricted license30 September 2026
19GLM 5.3Z.ai2.6%1.9%–3.3%Max#21 / 46#11$3.391MRestricted license30 September 2026
20GPT-6 LunaOpenAI1.7%0.3%–3.1%Max#22 / 46#31$0.501MAPI only30 September 2026
21Grok 4.6xAI1.2%0.1%–2.2%X-High#24 / 46—$6.00500KAPI only30 September 2026
22Grok 4.5xAI1.1%0.1%–2.1%Default#26 / 46—$6.00500KAPI only30 September 2026
23GPT-5.4OpenAI0.9%0.0%–1.8%High#27 / 46—$15.001MAPI only30 September 2026
24GPT-5.6 TerraOpenAI0.8%-0.4%–2.1%X-High#28 / 46#22$12.001MAPI only30 September 2026
25GLM 5.3 FlashZ.ai-0.3%-0.8%–0.2%Default#29 / 46#17$0.501MOpen weights30 September 2026
26GPT-5.6 LunaOpenAI-1.1%-1.8%–-0.4%X-High#31 / 46—$1.201MAPI only30 September 2026
27Gemini 3.7 FlashGoogle-1.6%-2.4%–-0.9%High#33 / 46—$3.751MAPI only30 September 2026
28Muse Spark 1.2Meta-3.0%-3.7%–-2.3%X-High#34 / 46—$4.251MAPI only30 September 2026
29Muse Spark 1.1Meta-4.6%-5.0%–-4.1%Default#35 / 46—$4.251MAPI only30 September 2026
30Qwen3.7 MaxQwen-5.2%-6.1%–-4.3%Default#36 / 46—$4.421MAPI only30 September 2026
31Hy3Tencent-5.7%-6.9%–-4.5%Default#37 / 46—$0.33262KOpen weights30 September 2026
32MiniMax M3MiniMax-6.7%-7.6%–-5.8%Default#38 / 46#40$1.201MRestricted license30 September 2026
33MiMo-V2.5-ProXiaomi-7.2%-8.1%–-6.3%Default#39 / 46—$0.871MOpen weights30 September 2026
34Qwen3.7 PlusQwen-7.3%-8.7%–-5.8%Default#40 / 46#30$1.281MAPI only30 September 2026
35Gemini 3.1 Pro PreviewGoogle-7.3%-8.3%–-6.4%Default#41 / 46#12$12.001MAPI only30 September 2026
36Gemini 3.6 FlashGoogle-7.7%-8.8%–-6.6%High#42 / 46—$3.751MAPI only30 September 2026
37Inkling SmallThinking Machines-9.8%-11.2%–-8.4%Default#43 / 46#82$1.20524KOpen weights30 September 2026
38InklingThinking Machines-10.9%-12.0%–-9.8%Default#44 / 46#47$4.05524KOpen weights30 September 2026
39Mistral Medium 3.5Mistral AI-11.7%-13.2%–-10.2%Default#45 / 46#65$7.50262KAPI only30 September 2026
40Solar Pro 4Upstage-15.8%-17.7%–-13.9%Default#46 / 46#119$0.36524KAPI only30 September 2026

How ModelCap uses LMArena Agent

LMArena Agent ranks models by agentic session score. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; LMArena Agent contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.

LMArena Agent leaderboard: common questions

What is the LMArena Agent benchmark?

LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. It is published by LMArena.

Which AI model leads LMArena Agent right now?

Claude Fable 5.1 (Anthropic) holds the top LMArena Agent score among the models ModelCap tracks, at 14.6% as of the source snapshot published 30 September 2026.

How many models are ranked on the LMArena Agent leaderboard here?

40 tracked models have a published LMArena Agent result on ModelCap; the source board itself lists 46 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on LMArena Agent?

Hy4 preview from Tencent is the highest-scoring model with openly downloadable weights on this board, at 4.2%.

Which model offers the best value on LMArena Agent?

Among the ten highest-scoring priced models, Hy4 preview has the lowest listed output price at $2.25 per 1M tokens while scoring 4.2%.

How is LMArena Agent scored?

The board ranks models by agentic session score; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does LMArena Agent decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; LMArena Agent contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the LMArena Agent results?

The newest LMArena Agent publication ModelCap holds is dated 30 September 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.

Explore further