Grading launch estimates: 12 settled records and the limits of the ledger
The captured ledger report now has 12 settled first predictions, a 9.3-point mean absolute error and 75% interval coverage. Reproducible outcome rows expose both the misses and the proof limits.
An estimated launch placement should be checked against what happens later. The useful object is the original prediction, with its interval and inputs, followed by a later qualifying outcome. Recalculating an estimate after seeing a result cannot validate forecasting.
We retained the live ModelCap report at the same cutoff as the model exports, inspected the first-prediction and settled-outcome ledger records read-only, and reproduced the full report with the existing producer. A minimal public extract of its 12 graded pairs allows the headline and channel arithmetic to be recomputed offline. The former draft's ten-outcome figures are superseded.
What qualifies under the inspected rule
The ledger grades revision zero, not a later revised estimate. A qualifying outcome arrives within 120 days, has a measured tier, includes a general board, and has at least two direct observations. Error is predicted minus measured: positive means the estimate was optimistic. The interval check uses the original bounds. Pairwise ordering checks outcome differences of at least five points.
For every retained pair, revision zero precedes the recorded outcome time, and the listed outcome boards include general evidence and at least two observations. That is an internal chronology check. It does not prove the upstream benchmark result was unavailable before the prediction: those first-public-availability receipts are outside this extract. Accordingly, we report these as the ledger's settled first predictions and its prospective label, with this proof limit visible.
The current report, reproduced
| Measure | Result |
|---|---|
| Graded first predictions | 12 |
| Mean absolute error | 9.3 points |
| Median signed error | −1.9 points |
| 80th-percentile absolute error | 13.9 points |
| Original intervals covering outcome | 9 / 12 (75%) |
| Correct eligible pair order | 30 / 46 (65.2%) |
Three original intervals missed. MiMo-V2.6-Flash predicted 84.2 with bounds 75.8–92.6 and settled at 73.8. Nemotron 3 Ultra predicted 70.8 with bounds 63.5–78.1 and settled at 60.7. GPT-6.1 Sol predicted 74.3 with bounds 60.7–87.9 and settled at 88.2. These are the ledger's outcome values at settlement, not the same models' scores in today's model export. A later board refresh must not rewrite the outcome against which the first prediction was graded.
Channel evidence remains small
| Channel | Outcomes | MAE | Signed median | p80 error | Coverage |
|---|---|---|---|---|---|
| Corpus prior | 4 | 13.9 | -10.5 | 21 | 100% |
| Family succession | 4 | 7.4 | -7.8 | 13.9 | 75% |
| Specialist measurement | 5 | 11.1 | -3.9 | 17.1 | 100% |
| Launch card | 4 | 6.5 | +6.5 | 10.4 | 50% |
A launch counts toward every channel with at least a quarter of its original blend, so the channel counts sum to 17, not 12. Each channel reports the blended prediction's error, not an isolated experiment on that term. The four card-associated outcomes have a +6.5 signed median and 50% coverage, which warrants attention but is too small to establish a general bias.
The inspected feedback threshold is eight outcomes per channel. None reaches it. The system therefore has no active ledger-based channel adjustment at this cutoff. The rule can apply an error-based uncertainty floor and a bounded optimistic-card haircut after the threshold; the threshold itself is an operational rule, not statistical proof of calibration.
Avoid three misleading interpretations
- Nine of twelve outcomes inside nominal 80% bounds cannot establish that the intervals are calibrated. One additional covered outcome would change coverage by 8.3 percentage points. The sample is small and selected by which launches became measured.
- Thirty of 46 pair comparisons are not 46 independent forecasting trials. Each launch appears in multiple pairs. The ratio is descriptive and does not by itself demonstrate a statistically significant advantage over chance.
- A near-zero median signed error can coexist with large misses in both directions. The 9.3-point absolute error is therefore worth reading beside the −1.9 signed median.
The report keeps four retrospective replay examples separate, and this article does not mix them into the 12 settled records. Models that never receive a qualifying outcome also cannot silently count as successes. The selected measured subset need not represent the entire launch population.
Make the audit reproducible
Download the settled outcome extract, captured report and offline analysis. The check recomputes absolute errors, original interval coverage, channel attribution and pair ordering. It also asserts strict ledger-time ordering and the outcome board rule.
For a stronger future audit, retain immutable first-prediction records and source-availability receipts before outcomes, report all eligible launches including missing outcomes, and compare publisher and lineage holdouts separately. More arithmetic agreement alone will not fill the missing upstream-time proof. Meanwhile, the estimated rows near the top should be read as candidates with disclosed uncertainty, not measured wins.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
- Captured launch report
September 30 first-party summary displayed on methodology and copied read-only from the live publication directory.
- Settled first-prediction extract
Twelve prediction/outcome pairs sufficient to recompute published statistics; ledger times do not prove first upstream availability.
Found a discrepancy? Report a correction with the article and source URL.