Which AI benchmarks does ModelCap track?
14 public boards: Arena text leaderboard; Arena coding leaderboard; LMArena Agent leaderboard; ARC-AGI-2 leaderboard; ARC-AGI-3 leaderboard; SWE-bench leaderboard; BFCL V4 function-calling leaderboard; Artificial Analysis Intelligence Index; Terminal-Bench 2.1 (Artificial Analysis run); τ²-Bench Telecom (Artificial Analysis run); Humanity's Last Exam (Artificial Analysis run); Terminal-Bench 4.0 leaderboard; Terminal-Bench 2.1 Terminus 2 leaderboard; WildClawBench OpenClaw leaderboard. Each has its own leaderboard page listing every tracked model with a published result.
Which AI model leads each benchmark right now?
Arena text: Claude Opus 4.6 (1505). Arena coding: Claude Fable 5 (1552). LMArena Agent: Claude Fable 5.1 (14.6%). ARC-AGI-2: GPT-6 Astra (95.0%). ARC-AGI-3: GPT-6 Astra (62.7%). SWE-bench: Gemini 3 Flash Preview (75.8%). BFCL V4: Claude Opus 4.5 (77.5). AA Intelligence Index: Claude Opus 5.5 (57.6). AA Terminal-Bench 2.1: Claude Fable 5.1 (91.4%). AA τ²-Bench Telecom: GLM 5.2 (99.1%). AA Humanity's Last Exam: Claude Opus 5.5 (61.4%). Terminal-Bench 4.0: GPT-6 Astra (58.2%). Terminal-Bench 2.1 Terminus 2: Claude Fable 5 (80.5). WildClawBench OpenClaw: GPT-5.6 Sol (67.2). Snapshot as of 2 October 2026.
What is the difference between Arena ratings and benchmark scores?
Arena ratings come from blind pairwise human votes converted to a Bradley–Terry (Elo-style) rating, so they measure preference. SWE-bench, ARC-AGI and BFCL are fixed task suites with a pass rate or accuracy score, so they measure a specific capability. Each leaderboard here shows its board's own units; the ModelCap Index puts every board on a common scale before combining them and never averages raw, incompatible scales.
How do these benchmarks feed the ModelCap Index?
When a model has results on the general boards (Arena's text leaderboard and the Artificial Analysis Intelligence Index), they anchor its ModelCap Index score, and its coding, agent and reasoning results move that anchor by fixed weights within a bounded band. Each result counts by its board's share of its family and by its confidence, which falls as the evidence ages, so a model measured on fewer boards carries less support and a wider interval. The methodology page documents the sources and weights.
Why is a model missing from a benchmark leaderboard here?
A model appears only when the source board publishes a result that matches a reviewed catalogue identity. Models the source has not evaluated, results that cannot be matched to an exact identity, and non-canonical variants are left off rather than guessed.
How often are the benchmark leaderboards updated?
Each source refreshes on its own cadence. Publication and page updates happen separately. Each row shows the publication date when the source provides one.