Skip to content
ModelCap

Benchmark leaderboard

LMArena Agent leaderboard

LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. This page lists every model ModelCap tracks with a published LMArena Agent result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 6 August 2026

Claude Opus 5 from Anthropic leads the LMArena Agent leaderboard among the 34 tracked models with 12.0% (High configuration), per the source snapshot published 6 August 2026.

Tracked models with a result
34
Entries on the source board
46
Open-weight models listed
5

LMArena Agent standings (34 models)

Sorted by published score, one row per model. Source: LMArena.

LMArena Agent leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Claude Opus 5Anthropic12.0%10.6%–13.4%High+1 more evaluated#1 / 46#2$25.001MAPI only6 August 2026
2Claude Fable 5Anthropic11.7%9.3%–14.0%High#2 / 46#1$50.001MAPI only6 August 2026
3GPT-5.6 SolOpenAI10.3%8.6%–12.0%X-High#4 / 46#7$30.001MAPI only6 August 2026
4Kimi K3Moonshot AI10.1%9.0%–11.1%Max#5 / 46#4$15.001MRestricted license6 August 2026
5Claude Opus 4.8Anthropic9.4%7.7%–11.0%Thinking+1 more evaluated#6 / 46$25.001MAPI only6 August 2026
6GPT-5.5OpenAI8.5%7.5%–9.4%X-High+2 more evaluated#7 / 46#10$30.001MAPI only6 August 2026
7Claude Opus 4.7Anthropic8.3%7.0%–9.6%Thinking+1 more evaluated#8 / 46$25.001MAPI only6 August 2026
8Claude Sonnet 5Anthropic7.5%5.3%–9.6%High#9 / 46#14$10.001MAPI only6 August 2026
9GLM 5.2Z.ai6.8%5.8%–7.7%Max#12 / 46#13$2.421MOpen weights6 August 2026
10Claude Opus 4.6Anthropic6.4%5.1%–7.7%Default#13 / 46$25.001MAPI only6 August 2026
11Grok 4.5xAI5.6%4.3%–6.8%Default#15 / 46#12$6.00500KAPI only6 August 2026
12GPT-5.4OpenAI5.3%4.4%–6.2%High#16 / 46$15.001MAPI only6 August 2026
13GPT-5.6 LunaOpenAI4.4%2.6%–6.2%X-High#17 / 46#19$0.601MAPI only6 August 2026
14GPT-5.6 TerraOpenAI3.1%1.8%–4.4%X-High#18 / 46#15$6.001MAPI only6 August 2026
15Claude Sonnet 4.6Anthropic3.0%1.7%–4.3%Default#20 / 46$15.001MAPI only6 August 2026
16Muse Spark 1.1Meta0.9%0.3%–1.5%Default#22 / 46$4.251MAPI only6 August 2026
17Kimi K2.7 CodeMoonshot AI0.7%-1.3%–2.6%Default#23 / 46#189$3.50262KRestricted license6 August 2026
18GLM 5.1Z.ai0.4%-0.4%–1.1%Default#24 / 46$4.40205KOpen weights6 August 2026
19Qwen3.7 MaxQwen-0.2%-1.0%–0.7%Default#25 / 46$4.421MAPI only6 August 2026
20Kimi K2.6Moonshot AI-0.5%-2.6%–1.7%Default#27 / 46$4.00262KRestricted license6 August 2026
21Gemini 3.1 Pro PreviewGoogle-0.5%-1.3%–0.2%Default#28 / 46#6$12.001MAPI only6 August 2026
22Gemini 3.5 FlashGoogle-0.7%-1.3%–-0.1%High+1 more evaluated#29 / 46$9.001MAPI only6 August 2026
23Hy3Tencent-1.1%-2.5%–0.3%Default#30 / 46#18$0.528262KOpen weights6 August 2026
24Qwen3.7 PlusQwen-1.9%-3.2%–-0.6%Default#31 / 46#16$1.281MAPI only6 August 2026
25MiniMax M3MiniMax-2.5%-3.3%–-1.6%Default#33 / 46#23$1.201MRestricted license6 August 2026
26Gemini 3.6 FlashGoogle-2.7%-4.1%–-1.4%Default#35 / 46#8$7.501MAPI only6 August 2026
27InklingThinking Machines-6.6%-7.5%–-5.7%Default#37 / 46#24$4.051MOpen weights6 August 2026
28Mistral Medium 3.5Mistral AI-6.8%-9.0%–-4.7%Default#38 / 46#30$7.50262KAPI only6 August 2026
29Grok 4.3xAI-8.3%-9.1%–-7.5%High+1 more evaluated#39 / 46$2.501MAPI only6 August 2026
30Grok Build 0.1xAI-8.8%-9.7%–-8.0%Default#40 / 46#212$2.00256KAPI only6 August 2026
31Gemini 3 Flash PreviewGoogle-8.9%-9.7%–-8.2%Default#41 / 46$3.001MAPI only6 August 2026
32Gemini 3.5 Flash LiteGoogle-10.4%-11.8%–-8.9%Default#42 / 46#17$2.501MAPI only6 August 2026
33Nemotron 3 UltraNVIDIA-13.7%-16.1%–-11.4%Default#44 / 46#220$3.60512KRestricted license6 August 2026
34Gemma 4 31BGoogle-17.6%-20.0%–-15.2%Default#46 / 46#20$0.34262KOpen weights6 August 2026

How ModelCap uses LMArena Agent

LMArena Agent ranks models by agentic session score. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as agent evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

LMArena Agent leaderboard: common questions

What is the LMArena Agent benchmark?

LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. It is published by LMArena.

Which AI model leads LMArena Agent right now?

Claude Opus 5 (Anthropic) holds the top LMArena Agent score among the models ModelCap tracks, at 12.0% as of the source snapshot published 6 August 2026.

How many models are ranked on the LMArena Agent leaderboard here?

34 tracked models have a published LMArena Agent result on ModelCap; the source board itself lists 46 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on LMArena Agent?

GLM 5.2 from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 6.8%.

Which model offers the best value on LMArena Agent?

Among the ten highest-scoring priced models, GLM 5.2 has the lowest listed output price at $2.42 per 1M tokens while scoring 6.8%.

How is LMArena Agent scored?

The board ranks models by agentic session score; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does LMArena Agent decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; LMArena Agent contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the LMArena Agent results?

The newest LMArena Agent publication ModelCap holds is dated 6 August 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Explore further