Four September launches, revisited with the evidence now available
A September 30 follow-up on Opus 5.5, GPT-6 Sol, Grok 4.7 and MiMo-V2.6-Pro. Broader evidence and product succession replace the old single-board story; missing launch-day inputs limit historical comparisons.
Claude Opus 5.5, GPT-6 Sol, Grok 4.7 and MiMo-V2.6-Pro carry catalogue release timestamps on September 21 and 22. A draft written around their arrival described a single-board launch race. The retained September 30 file now supports a more useful question: what evidence is available for those products today, and how should it affect a decision?
We do not publish the old launch-day ranks as verified history. The draft did not preserve a digest-identified September 23 input, and an all-version list today is not an archive of that day's ordering. Without the earlier file, subtracting an old draft score from a new score would give a number whose baseline is not independently checked.
The four rows at the retained cutoff
| Model | Catalogue release UTC | Current rank | All-version rank | Index | Support | Interval | Admitted boards |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 2026-09-22 16:32 | 2 | 2 | 88.4 | 65.4% | 79.9–96.9 | 5 |
| MiMo-V2.6-Pro | 2026-09-21 20:07 | 8 | 12 | 83.2 | 70.8% | 75.7–90.7 | 3 |
| Grok 4.7 | 2026-09-21 16:19 | 30 | 67 | 72.2 | 73.2% | 64.9–79.5 | 4 |
| GPT-6 Sol | 2026-09-22 18:12 | All versions only | 40 | 77.7 | 77.2% | 70.7–84.7 | 6 |
Catalogue release timestamps describe the catalogue's declared release field. They do not establish announcement time or the first instant a public endpoint was callable. Nor does a source row's publication date establish when ModelCap first ingested it. Preserve those time concepts separately when auditing a launch.
Which evidence is available now
Opus has Arena Overall, the AA Intelligence Index, Arena Coding, LMArena Agent and ARC-AGI-2. MiMo Pro has Arena Overall, AA and Arena Coding. Grok has those first three plus LMArena Agent. GPT-6 Sol has six observations, including both ARC tracks, but sits in the all-version scope while GPT-6.1 Sol occupies a current row.
The retained observations alone therefore do not support describing these four as single-general-board rows at publication. Broader coverage does not guarantee a higher point estimate: it can add disagreement. Grok's admitted normalized Arena Overall score is 69.4 and AA is 84.1. MiMo Pro's are 82.8 and 84.0, much closer. Source configurations remain disclosed; these are comparisons of admitted product evidence, not proof that every source ran an identical endpoint.
Understand the support arithmetic
One full-confidence AA result contributes 80 × 0.25 = 20 support points. A full-confidence Arena Overall result can add 80 × 0.75 = 60. Specialist observations add their own smaller contributions, scaled by confidence. This explains why a row with a high score can initially show 20% support and later broaden substantially.
At this cutoff, Opus has 65.4% support, MiMo Pro 70.8%, Grok 73.2% and archived Sol 77.2%. These figures describe current admitted coverage. They are not percentages of successful tasks and are not probabilities that a rank is correct.
The current leader, Sonnet 5.5, has only the AA observation and 20% support. Its leading position and wide 72.4–100 interval answer different questions: the score orders the point estimate; the interval exposes limited information. The broader-evidence price frontier excludes it under the measured 60% filter.
Succession is not score inflation
GPT-6 Sol's all-version rank is 40 at 77.7. GPT-6.1 Sol's current row is third at 88.2, with AA and both ARC tracks. That is what this cutoff says; it does not retrospectively establish GPT-6 Sol's launch placement or a task-level improvement between the products. Current-family representation and all-version ordering are separate catalogue scopes.
When comparing a current successor with an archived model, use the same snapshot, preserve both scopes, and inspect the configurations and shared benchmark rows. A newer version number cannot substitute for independent evidence. A single-board successor can also retain less coverage than its predecessor while occupying the current product seat.
What to retain for the next launch
- Save the model and evidence exports with generatedAt and input hashes at the time you cite a placement.
- Preserve the original estimate and interval separately from later score revisions and outcome records.
- Record the source publication time and the earliest retained receipt showing public availability; do not infer either from catalogue timestamps.
- Revisit the product when additional general or specialist evidence is admitted, explaining which inputs changed before interpreting the rank movement.
This follow-up replaces an unverifiable historical baseline with a checked current evidence map. It can help shortlist these products today. A true launch-to-follow-up score comparison still requires the missing preserved launch snapshot.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
Found a discrepancy? Report a correction with the article and source URL.