What is the LMArena Agent benchmark?
LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. It is published by LMArena.
Benchmark leaderboard
LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. This page lists every model ModelCap tracks with a published LMArena Agent result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 30 September 2026
Claude Fable 5.1 from Anthropic leads the LMArena Agent leaderboard among the 40 tracked models with 14.6% (Max configuration), per the source snapshot published 30 September 2026.
Sorted by published score, one row per model. Source: LMArena.
LMArena Agent ranks models by agentic session score. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; LMArena Agent contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.
LMArena's Agent board evaluates models running multi-step agentic sessions with tools; scores are the published win share with its official interval, per evaluated configuration. It is published by LMArena.
Claude Fable 5.1 (Anthropic) holds the top LMArena Agent score among the models ModelCap tracks, at 14.6% as of the source snapshot published 30 September 2026.
40 tracked models have a published LMArena Agent result on ModelCap; the source board itself lists 46 entries. Each row shows the model's best evaluated configuration.
Hy4 preview from Tencent is the highest-scoring model with openly downloadable weights on this board, at 4.2%.
Among the ten highest-scoring priced models, Hy4 preview has the lowest listed output price at $2.25 per 1M tokens while scoring 4.2%.
The board ranks models by agentic session score; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; LMArena Agent contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting.
The newest LMArena Agent publication ModelCap holds is dated 30 September 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.