Skip to content
ModelCap

Notes

What the board did, and why

The ranking is automatic; the reasons are not. These notes are written by the people who run ModelCap when a launch, an incident or a rule change is worth explaining in plain language, with the real numbers. They are not summaries of the table and there is no note per model.

  1. incidentssuccession

    The refresh that retired its own number one

    For about ninety minutes on 1 September 2026 the board hid Claude Fable 5, its measured #1, and opened Claude Fable 5.1 at #53 on a prior. Here is what the code did, why no alert fired, and the four rules that came out of it.

  2. guideindex

    How to read a ModelCap row

    Rank, score, point estimate, confidence, interval, evidence label and the movement arrow each mean one specific thing. This is the plain-language key, with the formula that ties them together and the mistakes people make most often.

  3. methodologyprinciples

    What the Index refuses to count

    Downloads, likes, prices, parameter counts, a publisher's own benchmark table, an LLM's opinion: each is available, each would make the board look more complete, and each is deliberately excluded from capability evidence. The reasons, and the one place a Hub error message forced us to be vaguer than we wanted.

  4. operationsautonomy

    How the board refreshes itself, and what happens when a refresh fails

    No human sits in the ranking loop and no scheduler lives in the source repository. A data plane discovers, admits, reads, scores, validates and publishes on a fixed cadence, keeps the last good board when a cycle fails, and reports its own health. The contract and the failures we have hit.