Skip to content
ModelCap

Notes

Four September launches, revisited with the evidence now available

By ModelCaplaunchesevidence

A September 30 follow-up on Opus 5.5, GPT-6 Sol, Grok 4.7 and MiMo-V2.6-Pro. Broader evidence and product succession replace the old single-board story; missing launch-day inputs limit historical comparisons.

Claude Opus 5.5, GPT-6 Sol, Grok 4.7 and MiMo-V2.6-Pro carry catalogue release timestamps on September 21 and 22. A draft written around their arrival described a single-board launch race. The retained September 30 file now supports a more useful question: what evidence is available for those products today, and how should it affect a decision?

We do not publish the old launch-day ranks as verified history. The draft did not preserve a digest-identified September 23 input, and an all-version list today is not an archive of that day's ordering. Without the earlier file, subtracting an old draft score from a new score would give a number whose baseline is not independently checked.

The four rows at the retained cutoff

The four rows at the retained cutoff
ModelCatalogue release UTCCurrent rankAll-version rankIndexSupportIntervalAdmitted boards
Claude Opus 5.52026-09-22 16:322288.465.4%79.9–96.95
MiMo-V2.6-Pro2026-09-21 20:0781283.270.8%75.7–90.73
Grok 4.72026-09-21 16:19306772.273.2%64.9–79.54
GPT-6 Sol2026-09-22 18:12All versions only4077.777.2%70.7–84.76

Catalogue release timestamps describe the catalogue's declared release field. They do not establish announcement time or the first instant a public endpoint was callable. Nor does a source row's publication date establish when ModelCap first ingested it. Preserve those time concepts separately when auditing a launch.

Which evidence is available now

Opus has Arena Overall, the AA Intelligence Index, Arena Coding, LMArena Agent and ARC-AGI-2. MiMo Pro has Arena Overall, AA and Arena Coding. Grok has those first three plus LMArena Agent. GPT-6 Sol has six observations, including both ARC tracks, but sits in the all-version scope while GPT-6.1 Sol occupies a current row.

The retained observations alone therefore do not support describing these four as single-general-board rows at publication. Broader coverage does not guarantee a higher point estimate: it can add disagreement. Grok's admitted normalized Arena Overall score is 69.4 and AA is 84.1. MiMo Pro's are 82.8 and 84.0, much closer. Source configurations remain disclosed; these are comparisons of admitted product evidence, not proof that every source ran an identical endpoint.

Understand the support arithmetic

One full-confidence AA result contributes 80 × 0.25 = 20 support points. A full-confidence Arena Overall result can add 80 × 0.75 = 60. Specialist observations add their own smaller contributions, scaled by confidence. This explains why a row with a high score can initially show 20% support and later broaden substantially.

At this cutoff, Opus has 65.4% support, MiMo Pro 70.8%, Grok 73.2% and archived Sol 77.2%. These figures describe current admitted coverage. They are not percentages of successful tasks and are not probabilities that a rank is correct.

The current leader, Sonnet 5.5, has only the AA observation and 20% support. Its leading position and wide 72.4–100 interval answer different questions: the score orders the point estimate; the interval exposes limited information. The broader-evidence price frontier excludes it under the measured 60% filter.

Succession is not score inflation

GPT-6 Sol's all-version rank is 40 at 77.7. GPT-6.1 Sol's current row is third at 88.2, with AA and both ARC tracks. That is what this cutoff says; it does not retrospectively establish GPT-6 Sol's launch placement or a task-level improvement between the products. Current-family representation and all-version ordering are separate catalogue scopes.

When comparing a current successor with an archived model, use the same snapshot, preserve both scopes, and inspect the configurations and shared benchmark rows. A newer version number cannot substitute for independent evidence. A single-board successor can also retain less coverage than its predecessor while occupying the current product seat.

What to retain for the next launch

  • Save the model and evidence exports with generatedAt and input hashes at the time you cite a placement.
  • Preserve the original estimate and interval separately from later score revisions and outcome records.
  • Record the source publication time and the earliest retained receipt showing public availability; do not infer either from catalogue timestamps.
  • Revisit the product when additional general or specialist evidence is admitted, explaining which inputs changed before interpreting the rank movement.

This follow-up replaces an unverifiable historical baseline with a checked current evidence map. It can help shortlist these products today. A true launch-to-follow-up score comparison still requires the missing preserved launch snapshot.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

Found a discrepancy? Report a correction with the article and source URL.

More notes