Open weights versus API only: access, evidence and a 5.2-point gap
The best open-labelled model with at least 60% measured support sits 5.2 Index points behind the comparable API-only leader. Access categories, primary licences and self-hosting costs need separate checks.
Choosing weights instead of a hosted product involves capability, rights and operations. The board provides a capability estimate and catalogue access metadata. It does not establish that a particular reader can download every listed checkpoint, legally use it for every purpose, or operate it economically.
We compare current models in the same retained snapshot. The conservative capability filter requires a measured, unfused score with at least 60% support; 78 models meet it. That is an evidence choice for this article, not a definition of production readiness.
Count the access categories accurately
| Export label | Current models | Measured, including blended | Priced models | Median listed blend $/M |
|---|---|---|---|---|
| gated | 20 | 14 | 10 | 0.135 |
| none | 39 | 38 | 39 | 1.900 |
| open | 153 | 59 | 53 | 0.310 |
| restricted | 64 | 20 | 13 | 0.370 |
| unknown | 3 | 0 | 1 | 0.425 |
Here none means no downloadable weights are linked in the source data. Open describes an open-licence classification. Gated requires an acceptance or access step. Restricted uses model-specific terms. Unknown means classification could not be verified. The 240 rows outside none are therefore not 240 proven, unrestricted self-hosting options. In particular, a repository's existence and an unverified licence are insufficient to assert rights.
Compare the strongest rows with broader evidence
| Access label | Model | Overall rank | Index | Support | Interval | Blended $/M |
|---|---|---|---|---|---|---|
| none | Claude Opus 5.5 | 2 | 88.4 | 65.4% | 79.9–96.9 | 8 |
| open | MiMo-V2.6-Pro | 8 | 83.2 | 70.8% | 75.7–90.7 | 0.54375 |
| restricted | Kimi K3 | 10 | 82 | 89.0% | 76.3–87.7 | 6 |
| gated | Gemma 3 27B | 132 | 44.5 | 62.7% | 39.5–49.5 | 0.1725 |
Claude Opus 5.5 leads this API-only subset at 88.4; MiMo-V2.6-Pro leads the open-labelled subset at 83.2. The 5.2-point difference is 88.4 − 83.2, not a percentage of tasks failed. Their intervals, 79.9–96.9 and 75.7–90.7, overlap. Kimi K3 leads the restricted-labelled subset at 82.0. These margins do not demonstrate settled differences in any single application.
The overall leader, Claude Sonnet 5.5 at 91.9, has 20% support and is excluded by the broader-evidence filter. If one compares it directly with MiMo, the nominal gap is 8.7 points, but that answers a weaker evidence question. State the filter before quoting a gap.
Check the repository behind the label
The exact repository associated with MiMo in the retained export is XiaomiMiMo/MiMo-V2.6-Pro-RL. A read-only check of its primary Hub metadata on September 30 reported licence mit, gated false, and revision 73875d00b30a89ef8cc353a0b60b0e9f9561952d. This establishes what that repository reported at inspection, not that every source benchmark evaluated that immutable revision.
The primary Kimi K3 licence and GLM 5.3 licence use model-specific terms. GLM 5.3 Flash's licence identifies MIT. These exact documents matter more than a blanket description of every downloadable model as open. Pin the revision, read the terms for your use and verify your own access before provisioning hardware.
Serving and measurement: association, not causation
Among the 153 open-labelled rows, 53 have a current endpoint and 46 of those are measured, including blended measurements. Of the 100 without an endpoint, 13 are measured. These are catalogue associations at one time. They do not establish that serving caused evaluation: popularity, age, model size and publication choices can affect both.
Endpoint counts describe routes observed by ModelCap, not guaranteed independent organizations or failover capacity. API-only median listed price is $1.90 per million blended tokens, compared with $0.31 for priced open-labelled models. Those cohort medians mix different capabilities and token economics; they are not a controlled open-versus-closed price experiment.
Self-hosting has a different cost axis
MiMo's $0.54375 is an API token listing, not your self-hosting cost. The latter includes accelerator capacity, utilization, power or cloud hours, serving software, storage, operations, queueing and redundancy. The feasibility of a quantization or local serving configuration also needs its own quality check. A benchmark on a published product does not automatically validate a compressed checkpoint on your hardware.
Use this analysis to find a few candidates with inspectable evidence. Then make the licence and deployment checks for those exact artifacts. The open-weight ranking has its own eligibility rule, so it need not include every restricted-labelled model discussed here. The price frontier article compares hosted token prices under explicit evidence filters.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
- MiMo primary repository
Exact repository from the retained export; public API metadata checked September 30 and recorded in the source-check artifact.
- Repository and source checks
Dated primary-source checks, immutable repository revision and explicit limits of licence and raw-interval inspection.
Found a discrepancy? Report a correction with the article and source URL.