The September board at a retained cutoff: 279 ranks and their evidence
A September 30 cutoff separates 131 measured-labelled rows from 148 modeled rows, records the top ten and source coverage, and identifies the evidence changes worth watching next month.
The September review uses a single saved cutoff rather than whichever live table happens to load while an article is read. At 20:54:24 UTC on September 30, ModelCap ranked 279 current language models and 418 across all versions. The extra 139 rows are all-version-only positions, not 139 preserved historical ranks.
The dataset contains a larger catalogue beyond those ranked rows. This article analyzes the current-ranked population only. Image and audio entries, unranked discoveries and superseded rows do not silently enter its denominators.
The top ten, with support visible
| Rank | Model | Catalogue release day | Index | Support | Admitted boards |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 | 2026-09-28 | 91.9 | 20.0% | 1 |
| 2 | Claude Opus 5.5 | 2026-09-22 | 88.4 | 65.4% | 5 |
| 3 | GPT-6.1 Sol | 2026-09-29 | 88.2 | 23.0% | 3 |
| 4 | Claude Fable 5.1 | 2026-09-01 | 86.9 | 83.2% | 5 |
| 5 | Muse Spark 1.3 | 2026-09-02 | 84.7 | 82.4% | 4 |
| 6 | GPT-6 Astra | 2026-09-04 | 83.6 | 81.1% | 6 |
| 7 | Qwen3.8 Max (0902) | 2026-09-03 | 83.3 | 20.0% | 1 |
| 8 | MiMo-V2.6-Pro | 2026-09-21 | 83.2 | 70.8% | 3 |
| 9 | Gemini 3.8 Flash | 2026-09-02 | 82.7 | 88.6% | 6 |
| 10 | Kimi K3 | 2026-07-16 | 82 | 89.0% | 6 |
Sonnet 5.5 leads at 91.9 with 20% support. Opus is second at 88.4 with 65.4%. Nine of these ten carry September catalogue release dates; Kimi K3 carries July. Release metadata is not proof of earliest announcement or endpoint availability. The table separates point estimate from evidence breadth instead of treating a top rank as a settled conclusion.
Measured and modeled positions
| Evidence state | Models | Median support | Median interval width | Top-25 rows |
|---|---|---|---|---|
| Measured | 124 | 62.75% | 11.40 | 20 |
| Measured · blended | 7 | 2.10% | 45.20 | 0 |
| Modeled · lineage | 90 | 17.05% | 33.80 | 0 |
| Modeled · peer benchmarks | 54 | 29.00% | 22.05 | 5 |
| Modeled · prior | 2 | 25.90% | 43.75 | 0 |
| Modeled · succession | 2 | 24.05% | 21.50 | 0 |
There are 131 measured-labelled rows, including seven specialist-only blended rows, and 148 modeled rows. A measured label is not necessarily broad evidence. Among the 124 unfused measured rows, 78 meet the article's 60% support threshold. Five of the top 25 remain modeled peer-benchmark placements. The launch audit checks original predictions rather than grading those changing current estimates after the fact.
Which observations count at this cutoff
| Board | Family share | Current rows | Newest admitted publication in export |
|---|---|---|---|
| Arena Overall | 0.75 | 105 | 2026-09-25 |
| Artificial Analysis Intelligence Index | 0.25 | 53 | 2026-09-30 |
| SWE-bench bash-only | 0.45 | 0 | No admitted rows |
| Arena Coding | 0.3 | 105 | 2026-09-25 |
| Terminal-Bench 2.1 Terminus 2 | 0.25 | 0 | No admitted rows |
| LMArena Agent | 0.5 | 21 | 2026-09-28 |
| BFCL V4 | 0.25 | 0 | No admitted rows |
| WildClawBench OpenClaw | 0.25 | 9 | 2026-07-20 |
| ARC-AGI-2 Verified | 0.7 | 25 | 2026-09-30 |
| ARC-AGI-3 Verified | 0.3 | 6 | 2026-09-30 |
These dates describe the admitted observations in the retained file. A live upstream board can publish newer results that have not entered the same admitted cohort. The absence of admitted rows does not mean a benchmark organization publishes no other tracks or versions. SWE-bench bash-only, Terminal-Bench 2.1 Terminus 2 and BFCL V4 currently supply no scored observations in this file.
With the registered shares and those three absent boards, the maximum available full-confidence support is 80 + 12 × 0.30 + 5 × 0.75 + 3 = 90.35%. That ceiling is a consequence of coverage, not a cap on capability. The coding article explains why all 105 admitted coding rows consequently come from one board.
Publisher counts without guessed aliases
| Reviewed publisher key | Current rows | Top 25 | Top 50 |
|---|---|---|---|
| qwen | 37 | 2 | 5 |
| nvidia | 24 | 0 | 0 |
| microsoft | 22 | 0 | 0 |
| 19 | 2 | 5 | |
| openai | 17 | 4 | 6 |
| allenai | 16 | 0 | 0 |
| ibm-granite | 11 | 0 | 0 |
| meta-llama | 8 | 1 | 2 |
The model export has 52 publisher display strings, while the underlying captured publication has 55 literal publisher keys. We obtained a narrow read-only projection of just model identity and publisher key from that exact generation, then applied the site's reviewed aliases and Hub release-authority mappings. This yields 45 reviewed publisher groups, including one Qwen group of 37 rows instead of separate Qwen/qwen keys. The retained identity map discloses every mapping and its reviewed authority source.
These are catalogue groups, not an independently audited count of legal entities. Alias authority records were reviewed on August 9 and can drift; this article applies them as the site's recorded policy rather than silently inferring company ownership from similar names. Catalogue counts reflect product breadth and ingestion as well as release activity; they do not measure publisher strength.
Price and context coverage
There are 116 rows with both listed token prices; their median 3:1 blend is $0.4437 per million combined tokens. The three-frontier article provides the full calculation under different evidence filters. A list price does not describe all-task cost or self-hosting economics.
Context length is listed for 121 current rows. The median is 262,144 tokens, 36 list at least one million, and the largest listed value is two million. These are catalogue limits, not measured long-context retrieval accuracy. Missing context values remain missing rather than being treated as small windows.
What deserves a new check in October
Watch for broader evidence on the high-ranked low-support rows, changes to the modeled top-25 cohort, and new admitted coding coverage. The newest admitted WildClawBench publication is July 20. It reaches 90 days on October 18; under the strictly older-than-90-day rule it becomes ineligible after the boundary unless a newer qualifying observation is published. If that source drops out alone, the full-confidence support ceiling falls by 1.25 points to 89.10%.
The next monthly article should use another retained cutoff and distinguish genuine new observations, configuration selection, product succession and recalibration. This September article's tables remain fixed. They support reproducible analysis of the board as captured; they do not predict launch dates, promise approvals, or certify deployment readiness for a model.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
Found a discrepancy? Report a correction with the article and source URL.