Score v9: three fixed-harness boards admitted, and a 30% cap on any one evaluator
The ninth revision of the ModelCap Score adds the Artificial Analysis runs of Terminal-Bench 2.1, Tau2-Bench Telecom and Humanity's Last Exam, one per specialist family, and asserts that no organisation other than the Arena anchor can carry more than 30% of any score.
The ModelCap Score is the point estimate underneath every rank: a weighted combination of a model's results across four evidence families, general, coding, agent and reasoning, each fed by named public boards with named shares. Version 9 shipped on 1 September 2026. It changes what the coding, agent and reasoning families are fed by, and it adds a structural limit the score has never had.
Why v9 exists
The day after Claude Fable 5.1 launched, the two products of that family sat a few tenths of a point apart on the board while the independent evaluations that had already measured both showed a clear gap: 65.7 against 62.1 on the Artificial Analysis Intelligence Index, 91.4 against 84.6 on a fixed-harness run of Terminal-Bench 2.1. The gap was invisible because the specialist families were scored only on boards that had not measured the new release, and in two cases had stopped publishing altogether. A frontier release was waiting on upstream boards that, for a closed model, sometimes never arrive. That story is told in the dormancy note.
Artificial Analysis measures every notable model on the same harness within days of release and publishes the per-evaluation results next to its index. Those results are third-party measurements of a model the publisher does not control, which is the class of evidence the score already admitted for the index itself. v9 admits three of them, one per specialist family.
Family sources after v9
| Family (weight) | Boards and shares |
|---|---|
| General (80) | Arena Overall 0.75 · Artificial Analysis Intelligence Index 0.25 |
| Coding (12) | SWE-bench bash-only 0.25 · Arena Coding 0.35 · AA Terminal-Bench 2.1 0.40 |
| Agent (5) | LMArena Agent 0.40 · BFCL v4 0.15 · WildClawBench OpenClaw 0.15 · AA Tau2-Bench Telecom 0.30 |
| Reasoning (3) | ARC-AGI-2 Verified 0.50 · ARC-AGI-3 Verified 0.20 · AA Humanity's Last Exam 0.30 |
Family weights are unchanged. The three new boards are percentage boards, pass rates from 0 to 100, so they take the same 80/20 split between placement and raw achievement that SWE-bench and ARC already use. Each feeds only the family it measures: a terminal pass rate says nothing about reasoning, and the scorer never lets it.
The organisation share cap
Admitting four Artificial Analysis boards raised a question the score had never had to answer: how much of a rank should one evaluator be allowed to carry? v9 answers it explicitly. Every entry in the source plan names the organisation that runs the board, an organisation's share is the sum over its boards of family weight times source share, and no organisation other than the Arena anchor may exceed 30%.
| Organisation | Share | Boards |
|---|---|---|
| Arena (anchor, exempt) | 0.662 | Arena Overall, Arena Coding, LMArena Agent |
| Artificial Analysis | 0.272 | Intelligence Index, Terminal-Bench 2.1, Tau2-Bench Telecom, Humanity's Last Exam |
| SWE-bench | 0.030 | SWE-bench bash-only |
| ARC Prize | 0.021 | ARC-AGI-2 Verified, ARC-AGI-3 Verified |
| BFCL | 0.0075 | BFCL v4 |
| WildClaw | 0.0075 | WildClawBench OpenClaw |
The cap is asserted when the scoring module loads. A plan entry without an organisation, or a non-anchor organisation above 0.30, throws before any score is computed. Raising an Artificial Analysis share is therefore a deliberate methodology change that cannot happen by accident, and one evaluator's outage or drift can move at most 30% of any score.
Also in v9
- Terminal-Bench 2.1 Terminus 2 is retired from scoring; the served Terminal-Bench 4.0 board is carried as display evidence with the vendor harness and effort named on each row, never scored.
- The Artificial Analysis evaluation boards date no individual row, so each model-and-board pair keeps the fetch instant that first showed it as its publication time, advancing only when the value or placement moves. A board 223 models wide is a field of 223, not the whole index.
- A board the payload no longer carries publishes as an unavailable source with zero rows; retained rows keep scoring until dormancy, and the lifecycle alerts make the loss visible long before that.
- The board shows the interval midpoint as “est.” beneath the score while a model's evidence is below strong support, so a floored or thin row is readable at a glance.
- Release authority stops demoting a lab's only checkpoint for being stored in FP8 or MXFP4; twelve first-party products got their board seats back. That one has its own note.
The full v9 text, including the measured effect on the board when it shipped, is the current score section of the methodology; v8 moved to the history folder of the repository.