Skip to content
ModelCap

Benchmark leaderboards

AI benchmark leaderboards

Every public benchmark board ModelCap ingests, as its publisher reports it: human-preference Arena ratings, agentic and function-calling suites, coding pass rates and ARC-AGI reasoning. Each board has its own leaderboard listing every tracked model with a published result, the source's rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 17 August 2026

Public boards tracked
7
Models with at least one result
199
Boards led by an Index-ranked model
5

Leaderboards

AI benchmarks: common questions

Which AI benchmarks does ModelCap track?

7 public boards: Arena coding leaderboard; LMArena Agent leaderboard; ARC-AGI-2 leaderboard; ARC-AGI-3 leaderboard; SWE-bench leaderboard; BFCL V4 function-calling leaderboard; Artificial Analysis Intelligence Index leaderboard. Each has its own leaderboard page listing every tracked model with a published result.

Which AI model leads each benchmark right now?

Arena coding: Claude Fable 5 (1554). LMArena Agent: Claude Opus 5 (12.2%). ARC-AGI-2: GPT-5.6 Sol (92.5%). ARC-AGI-3: Claude Opus 5 (30.2%). SWE-bench: Gemini 3 Flash Preview (75.8%). BFCL V4: GLM 4.6 (72.4). Artificial Analysis Intelligence Index: Claude Opus 5 (63.1). Snapshot as of 17 August 2026.

What is the difference between Arena ratings and benchmark scores?

Arena ratings come from blind pairwise human votes converted to a Bradley–Terry (Elo-style) rating, so they measure preference. SWE-bench, ARC-AGI and BFCL are fixed task suites with a pass rate or accuracy score, so they measure a specific capability. ModelCap shows each in its own units and never averages them into one number.

How do these benchmarks feed the ModelCap Index?

The ModelCap Index combines the public boards a model has an admitted result on, weighting by coverage and publishing an uncertainty interval; a model with a single board result carries a wider interval than one measured on several. The methodology page documents the sources and weights.

Why is a model missing from a benchmark leaderboard here?

A model appears only when the source board publishes a result that matches a reviewed catalogue identity. Models the source has not evaluated, results that cannot be matched to an exact identity, and non-canonical variants are left off rather than guessed.

How often are the benchmark leaderboards updated?

The data plane refreshes each source on its own cadence and seals the result into the dataset the site serves; every leaderboard page re-renders within a minute of a new dataset. Each row shows the publication date the source provides.

Explore further