Charts
AI model price vs performance, and the frontier over time
Two views of the ModelCap Index, drawn from the same numbers as the main board. The first sets every current ranked model that lists an API price against its Index score and traces the best-value line: the highest measured score available at or below each price. The second follows the best measured score by the date models were first listed.
Snapshot as of 23 September 2026
Capability vs price
115 of 274 current ranked models list an API price. The best-value line runs through 12 measured models, from gpt-oss-20b ($0.036 per million tokens, score 39.1) to Claude Opus 5.5 ($8.00, score 97.1).
Prices are the listed API prices from OpenRouter, blended three parts input to one part output, in US dollars per million tokens. The price axis is logarithmic: each gridline is ten times the one before. Inherited and estimated scores are modeled rather than measured, so they are drawn but never form the line. Models with no listed API price, including open-weight models you can only run yourself, are not drawn. Listed free of charge, and so off the logarithmic axis: Dots3-Note Preview (61.5) and North Mini Code (54.6).
The best-value line, cheapest first
The frontier over time
17 models have each set a new best measured score when listed, from GPT-4 (27.2, listed 28 May 2023) to Claude Opus 5.5 (97.1, listed 22 Sept 2026).
Each dot is a measured model on the all-versions board, placed at the date it was first listed in a public catalogue, usually OpenRouter, which can trail the publisher’s announcement. Scores are today’s Index scores, so the line shows how today’s evidence ranks models by when they appeared, not what was known at the time. Superseded models stay in.
Models that set a new high
How to read these charts
- Same numbers as the board. Every score is the published Index v4.12 score; nothing is adjusted for the charts.
- Scores carry ranges. The dots are point scores. Each model page shows its score range and the positions that range spans; how ranks and ranges work.
- Thin evidence is marked. A measured score that rests on few benchmarks reads Preliminary or Limited evidence in the tables; what these labels mean.
- List price is not the bill. Per-token list prices leave out caching and batch discounts, and a model that reasons at length spends more output tokens on the same task.
- Check the numbers yourself. The public dataset carries every score and listed price behind these charts.