How to read a ModelCap row
Rank, score, point estimate, confidence, interval, evidence label and the movement arrow each mean one specific thing. This is the plain-language key, with the formula that ties them together and the mistakes people make most often.
A row on the board packs six numbers and a label into one line. Most of the questions we get are not about the ranking method; they are about which of those numbers the reader should be looking at. This note is the key.
Rank
Rank orders current, canonical language models by their score, which is the lower end of the uncertainty interval. It is deliberately the pessimistic statistic. A model does not reach the top on one spectacular result; it reaches the top when enough independent boards agree that even a cautious reading puts it there. Every model page also shows an all-versions rank that includes superseded releases, so an older model that dropped off the current board still has a position among everything that ever ranked.
Score, point estimate, confidence, interval
| Number | What it is | Where it comes from |
|---|---|---|
| Score | Lower end of the interval; the ranking statistic | point estimate minus half the interval width |
| Point estimate | Best single estimate of capability | weighted, breadth-adjusted evidence across families |
| Confidence | Share of total scoring weight that has observed the model | sum of the weights of the boards it appears on |
| Interval | Range the board is prepared to stand behind | 4 + (100 − confidence) × 0.16 wide, centred on the point |
Two things follow from the formula. First, confidence is coverage, not certainty: a model at confidence 27 has been seen by boards carrying 27 of the 100 points of weight, and the biggest single weight is Arena Overall at 60, so a model Arena has not rated cannot get above about 40 whatever else it has done. Second, the interval is at least four points wide even at full coverage, so two models within a couple of points of each other are not meaningfully ordered by the board, and we would rather you read them as a tie.
The evidence label
- Measured. Independent public boards measured the exact catalogue product. It does not mean ModelCap ran the model; nothing on the site is our own benchmark.
- Measured · succession. Measured, but the score is a floor one tick above a measured predecessor because the successor's own interval is still wide. Read the point estimate for the actual estimate.
- Modeled · peer benchmarks. Placed from the publisher's own comparison table against models already on the board. The numbers are quarantined claims; the placement is a conservative reading of them.
- Modeled · lineage. Inherited from one verified base model with a penalty, larger for quantizations. A child can never duplicate its parent's score.
- Modeled · metadata. Estimated from nearby measured models by architecture, type, release era and size, shrunk toward the corpus median. Least reliable for novel architectures and metadata-poor closed models, and the label says so.
- Modeled · succession. A launch with no evidence and an older measured family member opens three points beneath it. It exists so a closed API launch does not open on a prior.
- Unranked. A written abstention: the identity is ambiguous, the card was unreadable, the peers could not be resolved, or the placement was uncalibrated. We show the gap rather than a guess.
The movement arrow
The arrow beside a rank is the change over the last 24 hours, measured between sealed boards. It moves for three reasons and the model page tells you which: new evidence for this model, new evidence for a neighbour, or a method change that re-scored everyone. A method change is stamped with its version on every row it touched, so a jump on the day of a release like version 4.8 is attributable.
Dates
Every fact on a model page carries the instant it is true as of. The board's own timestamp is when the snapshot was sealed, roughly every fifteen minutes. A benchmark observation carries the date its source published it; evidence older than a year is still measured but loses confidence and is labelled stale. Prices and provider uptime are as of the last catalogue read. We never substitute the request time for any of these, so a page you load tomorrow shows the same dates unless the underlying data changed.
What a model page adds
The row is the summary; the page is the audit trail. It lists every admitted observation with its board, configuration and publication date, the identity binding (exact catalogue product, pinned repository revision, or exact route), any lineage edge, the quantization if declared, and the versions of the scoring method that produced the current numbers. If you want to check a rank, the page is where the evidence is, and the public dataset carries the same fields for every model at once.