Skip to content
ModelCap
Language models274Newest top 10GPT-6 Sol#2 · 22 Sept 2026ModelCap #1Claude Opus 5.5Data updated 23 Sept 2026, 19:55 UTC
Recently available
More
Trending Local Models
More

AI model rankings — live LLM leaderboard

274 current Index positions

The full scored language board, including the specialised products listed under Other. Official publisher releases only; quantizations and repacks stay discoverable, unranked.

Current-model ModelCap rankings. Activate a column heading to change the sort order.
1Current-model rank
Anthropic
97.12 public benchmark observations across 2 boards; Index basis Measured; 22% support; score interval 72.4–100.0.
1M
128K max output
Exact context: 1,000,000 tokens; exact maximum output: 128,000 tokens
No Arena rating
$20.00
2Current-model rank
OpenAI
88.41 public benchmark observation across 1 board; Index basis Measured; 20% support; score interval 67.3–100.0.
1M
128K max output
Exact context: 1,050,000 tokens; exact maximum output: 128,000 tokens
No Arena rating
$10.00
3Current-model rank
87.85 public benchmark observations across 5 boards; Index basis Measured; 78% support; score interval 80.6–95.0.
1M
128K max output
Exact context: 1,000,000 tokens; exact maximum output: 128,000 tokens
1,498
LMArena
Highest supported configuration: Max; source rank 5
$50.00
4Current-model rank
xAI
87.41 public benchmark observation across 1 board; Index basis Measured; 20% support; score interval 66.3–100.0.
500K
450K max output
Exact context: 500,000 tokens; exact maximum output: 450,000 tokens
No Arena rating
$4.80
5Current-model rank
87.31 public benchmark observation across 1 board; Index basis Measured; 20% support; score interval 66.2–100.0.
1M
131K max output
Exact context: 1,048,576 tokens; exact maximum output: 131,072 tokens
No Arena rating
$0.87
6Current-model rank
86.51 public benchmark observation across 1 board; Index basis Measured; 20% support; score interval 65.4–100.0.
1M
131K max output
Exact context: 1,000,000 tokens; exact maximum output: 131,072 tokens
No Arena rating
$6.00
7Current-model rank
86.26 public benchmark observations across 6 boards; Index basis Measured; 68% support; score interval 77.6–94.8.
1M
128K max output
Exact context: 1,050,000 tokens; exact maximum output: 128,000 tokens
1,480
LMArena
Highest supported configuration: Max; source rank 24
$50.00
8Current-model rank
85.64 public benchmark observations across 4 boards; Index basis Measured; 75% support; score interval 78.0–93.2.
1M
944K max output
Exact context: 1,048,576 tokens; exact maximum output: 943,718 tokens
1,493
LMArena
Highest supported configuration: Max; source rank 8
$4.25
9Current-model rank
84.14 public benchmark observations across 4 boards; Index basis Measured; 75% support; score interval 76.5–91.7.
1M
66K max output
Exact context: 1,048,576 tokens; exact maximum output: 65,536 tokens
1,493
LMArena
Highest supported configuration: High; source rank 9
$3.75
10Current-model rank
83.9publisher-corpus-prior over 223 held-out anchors (8%); launch card against 5 resolved peers on 7 rows (92%); exceeds every named peer on 1 of 7 rows; no cross-lab optimism probe available; shrunk 0.6 toward the measured corpus; Index basis Estimated; 33% support; score interval 75.4–92.3.
1M
131K max output
Exact context: 1,048,576 tokens; exact maximum output: 131,072 tokens
No Arena rating
$0.28
11Current-model rank
83.84 public benchmark observations across 4 boards; Index basis Measured; 83% support; score interval 77.1–90.5.
1.3M
131K max output
Exact context: 1,310,720 tokens; exact maximum output: 131,072 tokens
1,483
LMArena
Highest supported configuration: Max; source rank 19
$2.64Increased by $0.59
12Current-model rank
83.56 public benchmark observations across 6 boards; Index basis Measured; 89% support; score interval 77.4–89.6.
1M
944K max output
Exact context: 1,048,576 tokens; exact maximum output: 943,718 tokens
1,485
LMArena
Highest supported configuration: Max; source rank 17
$15.00
13Current-model rank
OpenAI
82.27 public benchmark observations across 7 boards; Index basis Measured; 90% support; score interval 76.2–88.2.
1M
128K max output
Exact context: 1,050,000 tokens; exact maximum output: 128,000 tokens
1,482
LMArena
Highest supported configuration: High; source rank 20
$30.00
14Current-model rank
81.64 public benchmark observations across 4 boards; Index basis Measured; 82% support; score interval 74.9–88.3.
1.3M
944K max output
Exact context: 1,310,720 tokens; exact maximum output: 943,718 tokens
1,475
LMArena
Highest supported configuration: Default; source rank 29
$0.50
15Current-model rank
81.51 public benchmark observation across 1 board; Index basis Measured; 20% support; score interval 60.4–100.0.
1M
131K max output
Exact context: 1,000,000 tokens; exact maximum output: 131,072 tokens
No Arena rating
$0.47
16Current-model rank
81.4global-corpus-prior over 223 held-out anchors (6%); launch card against 4 resolved peers on 6 rows (94%); no cross-lab optimism probe available; shrunk 1.6 toward the measured corpus; Index basis Estimated; 32% support; score interval 73.9–88.8.
1M
262K max output
Exact context: 1,048,756 tokens; exact maximum output: 262,144 tokens
No Arena rating
$1.20
17Current-model rank
81.24 public benchmark observations across 4 boards; Index basis Measured; 86% support; score interval 74.9–87.5.
1M
66K max output
Exact context: 1,048,576 tokens; exact maximum output: 65,536 tokens
1,487
LMArena
Highest supported configuration: Default; source rank 15
$12.00
18Current-model rank
81.02 public benchmark observations across 2 boards; Index basis Measured; 62% support; score interval 76.0–86.0.
2M
1.8M max output
Exact context: 2,000,000 tokens; exact maximum output: 1,800,000 tokens
1,470
LMArena
Highest supported configuration: Beta · 0309; source rank 40
$2.50
19Current-model rank
80.22 public benchmark observations across 2 boards; Index basis Measured; 23% support; score interval 59.3–100.0.
1M
944K max output
Exact context: 1,048,576 tokens; exact maximum output: 943,718 tokens
No Arena rating
$0.50Decreased by $0.10
20Current-model rank
80.1global-corpus-prior over 223 held-out anchors (6%); launch card against 6 resolved peers on 5 rows (94%); optimism haircut 0 from cross-lab-probe; shrunk 1.6 toward the measured corpus; Index basis Estimated; 27% support; score interval 72.7–87.6.
— max output
Exact context: 0 tokens
No Arena rating
Self-hosted: weights only · no listed API price
21
Increased by 2
Current-model rank
ByteDance Seed
79.6global-corpus-prior over 223 held-out anchors (11%); launch card against 2 resolved peers on 3 rows (90%); exceeds every named peer on 1 of 3 rows; no cross-lab optimism probe available; shrunk 2.6 toward the measured corpus; Index basis Estimated; 23% support; score interval 69.9–89.2.
262K
236K max output
Exact context: 262,144 tokens; exact maximum output: 235,929 tokens
No Arena rating
$2.50
22
Decreased by 1
Current-model rank
79.56 public benchmark observations across 6 boards; Index basis Measured; 89% support; score interval 73.4–85.6.
1M
128K max output
Exact context: 1,050,000 tokens; exact maximum output: 128,000 tokens
1,466
LMArena
Highest supported configuration: X-High; source rank 45
$12.00
23
Decreased by 1
Current-model rank
79.5global-corpus-prior over 223 held-out anchors (8%); launch card against 8 resolved peers on 8 rows (92%); exceeds every named peer on 1 of 8 rows; optimism haircut 0 from cross-lab-probe; shrunk 1.9 toward the measured corpus; Index basis Estimated; 38% support; score interval 71.2–87.8.
262K
236K max output
Exact context: 262,144 tokens; exact maximum output: 235,929 tokens
No Arena rating
$0.25
24Current-model rank
79.2global-corpus-prior over 223 held-out anchors (10%); launch card against 2 resolved peers on 11 rows (90%); exceeds every named peer on 3 of 11 rows; optimism haircut 0 from cross-lab-probe; shrunk 2.4 toward the measured corpus; Index basis Estimated; 40% support; score interval 69.9–88.5.
— max output
Exact context: 0 tokens
No Arena rating
Self-hosted: weights only · no listed API price
25Current-model rank
78.53 public benchmark observations across 3 boards; Index basis Measured; 60% support; score interval 73.3–83.7.
1M
384K max output
Exact context: 1,048,576 tokens; exact maximum output: 384,000 tokens
1,463
LMArena
Highest supported configuration: High; source rank 50
$1.39Decreased by $0.59

Top AI models right now

The first ten positions on the ModelCap Index as of 23 September 2026, with the listed output price and context window beside each. The full board above is searchable and re-orders itself on every refresh.

  1. #1Claude Opus 5.5Anthropic · Index 97.1 · $20.00/1M out · 1M context
  2. #2GPT-6 SolOpenAI · Index 88.4 · $10.00/1M out · 1M context
  3. #3Claude Fable 5.1Anthropic · Index 87.8 · $50.00/1M out · 1M context
  4. #4Grok 4.7xAI · Index 87.4 · $4.80/1M out · 500K context
  5. #5MiMo-V2.6-ProXiaomi · Index 87.3 · $0.87/1M out · 1M context
  6. #6Qwen3.8 Max (0902)Qwen · Index 86.5 · $6.00/1M out · 1M context
  7. #7GPT-6 AstraOpenAI · Index 86.2 · $50.00/1M out · 1M context
  8. #8Muse Spark 1.3Meta · Index 85.6 · $4.25/1M out · 1M context
  9. #9Gemini 3.8 FlashGoogle · Index 84.1 · $3.75/1M out · 1M context
  10. #10MiMo-V2.6-FlashXiaomi · Index 83.9 · $0.28/1M out · 1M context

Rankings by what you need

Ranking lists

Benchmark leaderboards

Popular comparisons

Providers, publishers, models

What the ModelCap Index measures

Public evidence only. Every position comes from named public benchmark boards — Arena text and coding, LMArena Agent, ARC-AGI, SWE-bench, BFCL and others — matched to an exact model identity. Downloads, hype and provider counts never move a rank.

Uncertainty is published. Each score carries an interval and an evidence label (measured, inherited, estimated), so a one-point gap between neighbours reads as what it is. The methodology and the public dataset let anyone re-derive the board.

Rank, price and availability together. Listed API prices, context windows, weight access and every provider’s endpoint are read from the same snapshot as the rank and re-render within a minute of each refresh; the changes feed records every movement.

AI model rankings: common questions

What is the best AI model right now?

Claude Opus 5.5 from Anthropic holds ModelCap Index position #1 as of 23 September 2026 with a score of 97.1, ahead of GPT-6 Sol (#2) and Claude Fable 5.1 (#3). The Index combines public benchmark boards with published uncertainty, so a narrow gap between neighbours is shown as one; the ordering re-renders within a minute of every refresh.

How are the AI model rankings calculated?

Every current, canonical language model with matched public benchmark evidence receives a ModelCap Index score (0–100) with an interval. Evidence comes from named public boards — Arena text and coding, LMArena Agent, ARC-AGI, SWE-bench, BFCL and others — and each result counts by its board's share of its family and by its confidence, which falls as the evidence ages. A model not yet measured on those boards can hold a modeled position, labelled as such and shown with its confidence and interval. The methodology page publishes the weighting, the placement and identity rules and the change log. Market signals (downloads, provider counts) never move a rank.

What is the best LLM for coding?

Kimi K3 from Moonshot AI leads ModelCap's best-for-coding list as of 23 September 2026 (ModelCap Index #12). The list ranks models measured on at least two admitted coding boards (Arena Coding, SWE-bench bash-only, Terminal-Bench 2.1 with Terminus 2) by the Index's coding component, then lists models measured on a single board by that board's score, with the figures beside each row.

What is the cheapest LLM API?

Among ranked models with a listed price, Ling 3.0 Flash is the cheapest by blended token price as of 23 September 2026 ($0.063 per 1M output tokens), at ModelCap Index #78. The cheapest-LLMs list orders every priced model that way, and 14 ranked models currently have a free tier on at least one provider.

What is the best open-source (open-weight) LLM?

MiMo-V2.6-Pro from Xiaomi is the highest-ranked model with openly downloadable weights as of 23 September 2026, at ModelCap Index #5. The open-weight list orders every such model by Index position with its licence class and listed price; ModelCap labels weight availability, not whether the full training stack is open source.

Which AI model has the largest context window?

Grok 4.20 Multi-Agent lists the largest published context window among ranked models as of 23 September 2026: 2M tokens. The context-window list orders every ranked model by published limit; individual providers can serve less than the maximum.

How often are the rankings updated?

The dataset refreshes continuously and every page re-renders within a minute of a new snapshot; this one was published 23 September 2026. New models appear as soon as they pass automatic identity checks, and every rank or price movement is recorded on the changes feed.

How many AI models and providers does ModelCap track?

2262 canonical models across language, image, audio and video catalogues, 274 of them with a public ModelCap Index position, served by 79 inference providers observed through OpenRouter as of 23 September 2026. The models directory lists every one, A–Z by publisher.