Skip to content
ModelCap

Notes

How Claude Fable 5.1 took first place from Claude Fable 5 by a tenth of a point

By the ModelCap operatorslaunchesindex

On launch day the board put Claude Fable 5.1 at #1 on 86.3 and Claude Fable 5 at #2 on 86.2. The gap looks like a rounding error. It is not: it is a floor, and the estimate behind it is four and a half points wide.

Anthropic listed Claude Fable 5.1 on OpenRouter at 18:03 UTC on 1 September 2026. By the end of the day the ModelCap board showed it at #1 with a score of 86.3 and its predecessor, Claude Fable 5, at #2 with 86.2. Several people asked the obvious question: if the new model is better, why is the gap one tenth of a point? The short answer is that the published score is the low end of an uncertainty interval, and 86.3 is a floor, not an estimate. This note walks through the actual numbers.

The two rows as published

Claude Fable 5.1 and Claude Fable 5 on the board at the end of launch day
RowRankScore (lower bound)Point estimateConfidenceInterval
Claude Fable 5.1#186.393.127.286.3 to 100
Claude Fable 5#286.288.695.186.2 to 91.0

The board ranks on the lower bound of the interval, on purpose. A model earns a high position with breadth of evidence, not with one good run. Fable 5 has been measured on seven independent boards since June, including Arena, so its interval is narrow: about five points wide. Fable 5.1 had been out for a few hours. Four boards had it. Its interval is nearly fourteen points wide, and its raw lower bound came out at 85.3, a point beneath Fable 5.

What confidence actually measures

Confidence on ModelCap is not a probability. It is weight coverage: the share of the scoring plan's total weight that has actually observed this model. The interval width follows directly from it:

On launch day Fable 5.1 was present on the Artificial Analysis Intelligence Index (weight 20), the Artificial Analysis run of Terminal-Bench 2.1 (4.8), ARC-AGI-2 (1.5) and Humanity's Last Exam (0.9). That is 27.2 of the 100 points of weight, so confidence 27.2, so width 15.6, so a lower bound 7.8 points beneath a point estimate of 93.1. The missing 72.8 is mostly one source: Arena Overall alone carries 60 points of weight, and Arena had not listed the model yet. Fable 5 was added to every Arena board the day after its own launch; Opus 5 took about a week. Until that happens the interval stays wide however good the model is.

Why it is not simply #2

Before 1 September the board would have shown Fable 5.1 at #2 on 85.3 and moved it up when Arena arrived. That would have been the wrong answer for a reason that has nothing to do with confidence: on every board that had measured both models, the new one was ahead.

Independent boards that had measured both Claude Fable models on launch day
BoardFable 5.1Fable 5
Artificial Analysis Intelligence Index65.7 (#1)62.1 (#6)
Artificial Analysis Terminal-Bench 2.191.484.6
Humanity's Last Exam (AA run)59.155.5
GPQA Diamond (AA run)93.792.6
ARC-AGI-2 Verified90.089.2
Vals Index67.8766.04

Index version 4.8, built and deployed the same day, adds a rule for exactly this case: a measured successor whose lower bound sits beneath its measured predecessor is floored one tick above it, and the row is labelled Measured · succession. The floor is withdrawn the moment the boards that have measured both models, weighted the way the score weights them, favour the older release. On launch day four shared scored boards backed the succession and none opposed it, so the floor stood: 86.2 plus 0.1. The point estimate is never touched, which is why the row shows 93.1 above an 86.3. We wrote up the rule itself in a separate note.

So how far ahead is it really?

The honest answer is the point estimates: 93.1 against 88.6, a gap of about four and a half points on the ModelCap scale. That is consistent with the outside world. Artificial Analysis has the two models three and a half points apart on its index with Opus 5 in between; Vals has them just under two points apart; the terminal-work gap is nearly seven points. Anthropic's own launch table claims larger gaps on its in-house evaluations, mostly two to four points and up to twenty-eight on one science benchmark, but those numbers are the publisher's and are not admitted as evidence. Not everything went up: on the older ARC-AGI-1 set Fable 5.1 scored 97.5 against Fable 5's 98.5, and on one legal-agent evaluation it went backwards. The board does not score those, but they are a useful reminder that a successor is not better at everything.

What happens next

  • When Arena lists Claude Fable 5.1, its confidence rises to roughly 87 and its lower bound to roughly 90, at which point the succession floor is irrelevant and the label returns to plain Measured.
  • If Arena rates it beneath Fable 5 by enough to swing the weighted share of shared boards, the floor is withdrawn and the two rows re-order on their own evidence. That is the rule working, not failing.
  • Both rows stay on the board. Since the launch-day incident described in the next note, a successor only retires its predecessor once it is measured itself, and a measured predecessor is only hidden when its family has another current measured member.

The full rule text is in the methodology; the two rows are Claude Fable 5.1 and Claude Fable 5.

More notes