Rebuild the ModelCap Index yourself: a checked September example
The public notebook reproduces all 279 current and 418 all-version positions from a retained input. A worked score explains the arithmetic and the upstream facts a rebuild cannot validate.
The public evidence export makes the published Index arithmetic inspectable. We executed the captured notebook offline against the retained file, rather than fetching a different board during the check. All 418 published scores had rebuildable inputs, and every current and all-version rank was reproduced.
What the file supplies
There are 279 current positions and 139 additional all-version-only rows. Each evidence record includes score, basis, confidence, identity tie-break key, ranks, interval and admitted observations. An observation names its board, configuration, publication time, source URL, normalized score, confidence and family share. Fused estimates instead disclose means, standard deviations, weights and provenance. The method block supplies the family weights and 0.15-point comparison tolerance.
Raw upstream ratings, sample sizes and full normalization cohorts are outside this export. A benchmark source link remains necessary when checking the measurement itself. The normalized scores here are ModelCap's transformations of source evidence.
A measured worked example
Claude Fable 5.1 is fourth at 86.9 with 83.2% support. Its exported observations are:
| Family | Board | Score | Share | Confidence | Configuration |
|---|---|---|---|---|---|
| general | Arena Overall | 87.8 | 0.75 | 94.3% | Max |
| general | Artificial Analysis Intelligence Index | 89.8 | 0.25 | 100.0% | intelligence-index:v4.3:reasoning-unspecified |
| coding | Arena Coding | 81.8 | 0.3 | 61.9% | Max |
| agent | LMArena Agent | 79.8 | 0.5 | 92.5% | Max |
| reasoning | ARC-AGI-2 Verified | 80.8 | 0.7 | 100.0% | Max |
The general score uses share × confidence as its weight. From the published rounded inputs, (87.8 × 0.75 × 0.943 + 89.8 × 0.25) ÷ (0.75 × 0.943 + 0.25) = 88.3223. Each specialist moves that general score by its fixed family weight times the difference, capped at ±25 points before weighting.
Coding contributes 0.12 × (81.8 − 88.3223) = −0.7827. Agent contributes 0.05 × (79.8 − 88.3223) = −0.4261. Reasoning contributes 0.03 × (80.8 − 88.3223) = −0.2257. The rebuilt score is 86.8878, matching the published 86.9 within tolerance.
Support follows a different sum: 80 × (0.75 × 0.943 + 0.25) + 12 × 0.30 × 0.619 + 5 × 0.50 × 0.925 + 3 × 0.70 = 83.2209. The producer retains more precise inputs than the one-decimal export. Across all 230 unfused measured rows, the maximum support reconstruction error was 0.0732 points.
The cap, ties and estimates
Gemini 3.1 Pro Preview's rounded inputs produce general 80.775 and agent 31.1. The agent difference is −49.675, but only −25 counts, so its agent step is −1.25. Across current unfused measured rows, 24 specialist differences in 22 models exceed the cap; all are negative in this cutoff.
Ranks sort on the published one-decimal score, then basis (measured before inherited before estimated), support and Unicode identity key. One current tie is GPT-5.5 and Grok 4.20 Multi-Agent, both at 79.8, in positions 13 and 14. The tie-break rules, rather than an undisclosed extra decimal, decide the order.
A row without general evidence can combine estimate terms by inverse variance: sum(mean ÷ σ²) ÷ sum(1 ÷ σ²). LongCat 2.0 combines a card term of 80.8 with σ 6.3 and a corpus prior of 56.3 with σ 23.0. That arithmetic yields about 79.09, published as 79.1. The card dominates because it has smaller stated uncertainty. This is an estimated placement, not a new independently measured result. The retained notebook checks the fused arithmetic too.
The actual offline check
| Check | Matched |
|---|---|
| Measured scores | 230 / 230 |
| Family scores | 562 / 562 |
| Fused scores | 188 / 188 |
| Specialist terms inside fusions | 11 / 11 |
| Current ranks | 279 / 279 |
| All-version ranks | 418 / 418 |
| Scores without rebuildable inputs | 0 |
Download the captured notebook and offline runner beside the evidence file. The runner sets MODELCAP_EVIDENCE to that local file and fails if a score or rank check misses. The retained output records this run.
What passing does and does not establish
Passing establishes arithmetic consistency between the inputs and published ordering. It does not establish that a parser selected the right checkpoint, that an evaluation used the same effort as another source, that a reported licence is correct, or that normalization is the best way to compare tasks. Those checks require the original source and the admission policy.
Save the snapshot when citing a result. Peer fields and fitted mappings can change between refreshes even when a particular model's raw benchmark result is unchanged. A current all-version rank is also a freshly computed ordering, not that model's rank on its release day. This article makes no unsupported September 23 movement claim.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
- Offline notebook output
Captured run against the retained September 30 evidence file; all checks passed within the public tolerance.
Found a discrepancy? Report a correction with the article and source URL.