Skip to content
ModelCap

Notes

The September board at a retained cutoff: 279 ranks and their evidence

By ModelCapanalysismonthly

A September 30 cutoff separates 131 measured-labelled rows from 148 modeled rows, records the top ten and source coverage, and identifies the evidence changes worth watching next month.

The September review uses a single saved cutoff rather than whichever live table happens to load while an article is read. At 20:54:24 UTC on September 30, ModelCap ranked 279 current language models and 418 across all versions. The extra 139 rows are all-version-only positions, not 139 preserved historical ranks.

The dataset contains a larger catalogue beyond those ranked rows. This article analyzes the current-ranked population only. Image and audio entries, unranked discoveries and superseded rows do not silently enter its denominators.

The top ten, with support visible

The top ten, with support visible
RankModelCatalogue release dayIndexSupportAdmitted boards
1Claude Sonnet 5.52026-09-2891.920.0%1
2Claude Opus 5.52026-09-2288.465.4%5
3GPT-6.1 Sol2026-09-2988.223.0%3
4Claude Fable 5.12026-09-0186.983.2%5
5Muse Spark 1.32026-09-0284.782.4%4
6GPT-6 Astra2026-09-0483.681.1%6
7Qwen3.8 Max (0902)2026-09-0383.320.0%1
8MiMo-V2.6-Pro2026-09-2183.270.8%3
9Gemini 3.8 Flash2026-09-0282.788.6%6
10Kimi K32026-07-168289.0%6

Sonnet 5.5 leads at 91.9 with 20% support. Opus is second at 88.4 with 65.4%. Nine of these ten carry September catalogue release dates; Kimi K3 carries July. Release metadata is not proof of earliest announcement or endpoint availability. The table separates point estimate from evidence breadth instead of treating a top rank as a settled conclusion.

Measured and modeled positions

Measured and modeled positions
Evidence stateModelsMedian supportMedian interval widthTop-25 rows
Measured12462.75%11.4020
Measured · blended72.10%45.200
Modeled · lineage9017.05%33.800
Modeled · peer benchmarks5429.00%22.055
Modeled · prior225.90%43.750
Modeled · succession224.05%21.500

There are 131 measured-labelled rows, including seven specialist-only blended rows, and 148 modeled rows. A measured label is not necessarily broad evidence. Among the 124 unfused measured rows, 78 meet the article's 60% support threshold. Five of the top 25 remain modeled peer-benchmark placements. The launch audit checks original predictions rather than grading those changing current estimates after the fact.

Which observations count at this cutoff

Which observations count at this cutoff
BoardFamily shareCurrent rowsNewest admitted publication in export
Arena Overall0.751052026-09-25
Artificial Analysis Intelligence Index0.25532026-09-30
SWE-bench bash-only0.450No admitted rows
Arena Coding0.31052026-09-25
Terminal-Bench 2.1 Terminus 20.250No admitted rows
LMArena Agent0.5212026-09-28
BFCL V40.250No admitted rows
WildClawBench OpenClaw0.2592026-07-20
ARC-AGI-2 Verified0.7252026-09-30
ARC-AGI-3 Verified0.362026-09-30

These dates describe the admitted observations in the retained file. A live upstream board can publish newer results that have not entered the same admitted cohort. The absence of admitted rows does not mean a benchmark organization publishes no other tracks or versions. SWE-bench bash-only, Terminal-Bench 2.1 Terminus 2 and BFCL V4 currently supply no scored observations in this file.

With the registered shares and those three absent boards, the maximum available full-confidence support is 80 + 12 × 0.30 + 5 × 0.75 + 3 = 90.35%. That ceiling is a consequence of coverage, not a cap on capability. The coding article explains why all 105 admitted coding rows consequently come from one board.

Publisher counts without guessed aliases

Publisher counts without guessed aliases
Reviewed publisher keyCurrent rowsTop 25Top 50
qwen3725
nvidia2400
microsoft2200
google1925
openai1746
allenai1600
ibm-granite1100
meta-llama812

The model export has 52 publisher display strings, while the underlying captured publication has 55 literal publisher keys. We obtained a narrow read-only projection of just model identity and publisher key from that exact generation, then applied the site's reviewed aliases and Hub release-authority mappings. This yields 45 reviewed publisher groups, including one Qwen group of 37 rows instead of separate Qwen/qwen keys. The retained identity map discloses every mapping and its reviewed authority source.

These are catalogue groups, not an independently audited count of legal entities. Alias authority records were reviewed on August 9 and can drift; this article applies them as the site's recorded policy rather than silently inferring company ownership from similar names. Catalogue counts reflect product breadth and ingestion as well as release activity; they do not measure publisher strength.

Price and context coverage

There are 116 rows with both listed token prices; their median 3:1 blend is $0.4437 per million combined tokens. The three-frontier article provides the full calculation under different evidence filters. A list price does not describe all-task cost or self-hosting economics.

Context length is listed for 121 current rows. The median is 262,144 tokens, 36 list at least one million, and the largest listed value is two million. These are catalogue limits, not measured long-context retrieval accuracy. Missing context values remain missing rather than being treated as small windows.

What deserves a new check in October

Watch for broader evidence on the high-ranked low-support rows, changes to the modeled top-25 cohort, and new admitted coding coverage. The newest admitted WildClawBench publication is July 20. It reaches 90 days on October 18; under the strictly older-than-90-day rule it becomes ineligible after the boundary unless a newer qualifying observation is published. If that source drops out alone, the full-confidence support ceiling falls by 1.25 points to 89.10%.

The next monthly article should use another retained cutoff and distinguish genuine new observations, configuration selection, product succession and recalibration. This September article's tables remain fixed. They support reproducible analysis of the board as captured; they do not predict launch dates, promise approvals, or certify deployment readiness for a model.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

Found a discrepancy? Report a correction with the article and source URL.

More notes