The ranking is automatic; the reasons are not. These notes are written by the people who run ModelCap when a launch, an incident or a rule change is worth explaining in plain language, with the real numbers. They are not summaries of the table and there is no note per model.
On launch day the board put Claude Fable 5.1 at #1 on 86.3 and Claude Fable 5 at #2 on 86.2. The gap looks like a rounding error. It is not: it is a floor, and the estimate behind it is four and a half points wide.
For about ninety minutes on 1 September 2026 the board hid Claude Fable 5, its measured #1, and opened Claude Fable 5.1 at #53 on a prior. Here is what the code did, why no alert fired, and the four rules that came out of it.
Index version 4.8 floors a measured successor one tick above its measured predecessor until the boards that have scored both say otherwise. It moved DeepSeek V4 Pro from #102 to #20 and settled a Fable question the same day. The rule, the traps, and what it did to the board.
Every rankable model opens on the board the cycle it is admitted. For open-weight releases the opening position comes from the publisher's own comparison table, read deterministically and quarantined. A worked example with Qwen3.8-27B, and the bug that nearly sent it to the corpus prior.
Rank, score, point estimate, confidence, interval, evidence label and the movement arrow each mean one specific thing. This is the plain-language key, with the formula that ties them together and the mistakes people make most often.
Three of the coding and agent boards the Index scored have stopped publishing rows the board can use. The dormancy rule that retires them, the Terminal-Bench 4.0 change that broke the fetch, and why vendor-harness results are shown but never scored.
The ninth revision of the ModelCap Score adds the Artificial Analysis runs of Terminal-Bench 2.1, Tau2-Bench Telecom and Humanity's Last Exam, one per specialist family, and asserts that no organisation other than the Arena anchor can carry more than 30% of any score.
Downloads, likes, prices, parameter counts, a publisher's own benchmark table, an LLM's opinion: each is available, each would make the board look more complete, and each is deliberately excluded from capability evidence. The reasons, and the one place a Hub error message forced us to be vaguer than we wanted.
A product should hold exactly one seat on the board however many ways it ships. Two rules get that right for API-plus-checkpoint twins. A third rule got it wrong for labs that release their only checkpoint in FP8, and quietly unseated GLM-5.3, DeepSeek V4 and gpt-oss until 1 September.
No human sits in the ranking loop and no scheduler lives in the source repository. A data plane discovers, admits, reads, scores, validates and publishes on a fixed cadence, keeps the last good board when a cycle fails, and reports its own health. The contract and the failures we have hit.