Skip to content
ModelCap

Notes

What capability costs: three price frontiers at the September cutoff

By ModelCappricinganalysis

A retained September 30 snapshot shows how an evidence filter changes the price frontier. Exact token prices, workload arithmetic and an offline recipe make the trade-offs inspectable.

A price frontier asks a specific question: which listed models offer more Index score than every cheaper alternative? It is a way to narrow a shortlist, not a claim that one model is cheapest for every task. On this cutoff, 116 of 279 current-ranked models have both token prices. The other 163 are absent from the calculation; a missing price is not zero.

The price and evidence rules

We use dollars per million tokens and a 3:1 input-to-output workload: 0.75 × input price + 0.25 × output price. The score is the published ModelCap Index, not a task pass rate. The retained models export supplies the exact price values used below; display rounding never changes the frontier membership.

Sort by price ascending, then score descending at equal prices. Keep a row only when its score exceeds the best already seen. That rule removes a same-price lower-scoring alternative. Apply the evidence filter before sorting, rather than removing inconvenient estimates afterwards.

Every priced model

Every priced model
Blended $/MModelIndexSupportEvidence
0Dots3-Note Preview60.12.1%Measured · blended
0.03115Ling 3.0 Flash VL66.320.0%Measured
0.04375Nex-N2.5-Mini71.423.0%Modeled · peer benchmarks
0.11385DeepSeek V4.1 Flash79.978.2%Measured
0.54375MiMo-V2.6-Pro83.270.8%Measured
2Muse Spark 1.384.782.4%Measured
4Claude Sonnet 5.591.920.0%Measured

The zero-price Dots3-Note Preview listing is retained because zero is the published value. It carries only 2.1% support and a specialist result blended with estimates. That row does not establish an unrestricted, permanently free service. Ling 3.0 Flash VL has one general result, Nex-N2.5-Mini is modeled, and the leading Sonnet has one general result. The nominal frontier therefore mixes very different kinds of evidence.

At least 25% published support

At least 25% published support
Blended $/MModelIndexSupportEvidence
0.036gpt-oss-20b37.980.3%Measured
0.07025gpt-oss-120b44.783.2%Measured
0.08438Qwen3 30B A3B Instruct 250750.663.0%Measured
0.11385DeepSeek V4.1 Flash79.978.2%Measured
0.54375MiMo-V2.6-Pro83.270.8%Measured
2Muse Spark 1.384.782.4%Measured
8Claude Opus 5.588.465.4%Measured

The leading Sonnet drops out and Opus enters. Support is the exported confidence field. On a measured row it describes admitted evidence mass. On a modeled row it is an estimate-support disclosure, not the fraction of benchmark families measured. A threshold alone is therefore insufficient as a guarantee of independent coverage.

Measured, without blended estimates, and at least 60% support

Measured, without blended estimates, and at least 60% support
Blended $/MModelIndexSupportEvidence
0.036gpt-oss-20b37.980.3%Measured
0.07025gpt-oss-120b44.783.2%Measured
0.08438Qwen3 30B A3B Instruct 250750.663.0%Measured
0.11385DeepSeek V4.1 Flash79.978.2%Measured
0.54375MiMo-V2.6-Pro83.270.8%Measured
2Muse Spark 1.384.782.4%Measured
8Claude Opus 5.588.465.4%Measured

There are 78 current models meeting this evidence rule, whether priced or not. In this particular snapshot its priced frontier equals the 25% frontier: no modeled candidate survives onto that frontier. The equality is a finding for this cutoff, not an equivalence between the filters. These models are still shortlisting evidence, not production certifications.

DeepSeek V4.1 Flash to MiMo-V2.6-Pro buys 3.3 Index points for about 4.78 times the blended list price (0.54375 ÷ 0.11385). MiMo to Muse Spark buys 1.5 points for about 3.68 times the price. Muse to Opus buys 3.7 points for four times the price. These are point-estimate differences; their score intervals overlap and their task value has not been measured here.

Turn the blend into your workload

For an illustrative run of 3 million input tokens and 1 million output tokens, MiMo's listed prices imply 3 × $0.435 + 1 × $0.87 = $2.175. DeepSeek's imply 3 × $0.0198 + 1 × $0.396 = $0.4554. These are arithmetic examples, not bills: caching, routing, retries, tool calls, reasoning tokens and provider terms can change actual cost.

Use the workload cost guide before applying the 3:1 list to output-heavy work. A model that produces fewer tokens or needs fewer retries can cost less per successful task despite a higher token price. The site's Best value method uses Arena quality and vote eligibility, so its shortlist can differ from this Index frontier.

Reproduce and use the result

The offline analysis walks the retained exports with the rule above and produces all three exact frontiers. Check its hashes, run it, and substitute your own token mixture if you want a different cost axis. Preserve the evidence filter and list the excluded models. Then evaluate the finalists on your own workload with the same configuration and a budget you choose; this article used no model inference.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

Found a discrepancy? Report a correction with the article and source URL.

More notes