Skip to content
ModelCap

Benchmark leaderboard

ARC-AGI-2 leaderboard

ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. This page lists every model ModelCap tracks with a published ARC-AGI-2 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 1 October 2026

GPT-6 Astra from OpenAI leads the ARC-AGI-2 leaderboard among the 67 tracked models with 95.0% (Max configuration), per the source snapshot published 1 October 2026.

Tracked models with a result
67
Entries on the source board
246
Open-weight models listed
10

ARC-AGI-2 standings (67 models)

Sorted by published score, one row per model. Source: ARC Prize Foundation.

ARC-AGI-2 leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1GPT-6 AstraOpenAI95.0%Max+5 more evaluated#1 / 246#8$50.001MAPI only1 October 2026
2GPT-6.1 SolOpenAI94.2%Max+4 more evaluated#2 / 246#3$10.001MAPI only1 October 2026
3Claude Opus 5.5Anthropic93.3%High+4 more evaluated#3 / 246#2$20.001MAPI only1 October 2026
4GPT-5.6 SolOpenAI92.5%Max+4 more evaluated#5 / 246—$10.001MAPI only1 October 2026
5Claude Opus 5Anthropic90.4%Max+1 more evaluated#12 / 246—$25.001MAPI only1 October 2026
6Claude Fable 5.1Anthropic90.0%Max+4 more evaluated#13 / 246#4$50.001MAPI only1 October 2026
7GPT-6 SolOpenAI89.6%Max+5 more evaluated#16 / 246—$10.001MAPI only1 October 2026
8Claude Fable 5Anthropic89.2%Max+4 more evaluated#17 / 246—$50.001MAPI only1 October 2026
9Gemini 3.8 FlashGoogle89.2%High+2 more evaluated#17 / 246#9$3.751MAPI only1 October 2026
10GPT-5.5OpenAI85.0%X-High+3 more evaluated#28 / 246#14$30.001MAPI only1 October 2026
11Gemini 3.7 FlashGoogle84.6%High+2 more evaluated#29 / 246—$3.751MAPI only1 October 2026
12GPT-5.5 ProOpenAI84.6%High+1 more evaluated#30 / 246#62$1801MAPI only1 October 2026
13GPT-5.6 TerraOpenAI83.9%Max+4 more evaluated#33 / 246#22$12.001MAPI only1 October 2026
14GPT-5.4 ProOpenAI83.3%X-High#34 / 246—$1801MAPI only1 October 2026
15Dots3-Note PreviewDots Studio76.8%Max#42 / 246#73Free512KAPI only1 October 2026
16GPT-5.4OpenAI74.0%X-High+3 more evaluated#47 / 246—$15.001MAPI only1 October 2026
17Claude Opus 4.8Anthropic72.1%High+2 more evaluated#48 / 246—$25.001MAPI only1 October 2026
18Gemini 3.5 FlashGoogle72.1%High+1 more evaluated#48 / 246—$9.001MAPI only1 October 2026
19Claude Opus 4.6Anthropic69.2%High+3 more evaluated#53 / 246—$25.001MAPI only1 October 2026
20Grok 4.6xAI67.1%X-High+3 more evaluated#59 / 246—$6.00500KAPI only1 October 2026
21GLM 5.3 FlashZ.ai65.8%Max+2 more evaluated#63 / 246#17$0.501MOpen weights1 October 2026
22Grok 4.20xAI65.1%Reasoning#64 / 246—$2.502MAPI only1 October 2026
23DeepSeek V4 Flash 0731DeepSeek61.4%Max+3 more evaluated#70 / 246—$1.281MOpen weights1 October 2026
24DeepSeek V4 Pro 0813DeepSeek61.3%Max+3 more evaluated#71 / 246#24$1.981MOpen weights1 October 2026
25Claude Sonnet 4.6Anthropic60.4%High+1 more evaluated#73 / 246—$15.001MAPI only1 October 2026
26Gemini 3.6 FlashGoogle60.4%High+3 more evaluated#74 / 246—$3.751MAPI only1 October 2026
27Kimi K3Moonshot AI60.4%Max+2 more evaluated#74 / 246#10$10.001MRestricted license1 October 2026
28GPT-5.6 LunaOpenAI59.5%Max+4 more evaluated#79 / 246—$1.201MAPI only1 October 2026
29GPT-6 LunaOpenAI59.3%Max+5 more evaluated#80 / 246#31$0.501MAPI only1 October 2026
30GPT-5.2 ProOpenAI54.2%High+1 more evaluated#87 / 246—$168400KAPI only1 October 2026
31GPT-5.2OpenAI52.9%X-High+4 more evaluated#89 / 246—$14.00400KAPI only1 October 2026
32Grok 4.5xAI52.6%High+2 more evaluated#90 / 246—$6.00500KAPI only1 October 2026
33Qwen3.8 27BQwen42.4%X-High+3 more evaluated#100 / 246#39$3.001MOpen weights1 October 2026
34Inkling SmallThinking Machines40.1%X-High+5 more evaluated#102 / 246#82$1.20524KOpen weights1 October 2026
35Claude Opus 4.5Anthropic37.6%Thinking+3 more evaluated#104 / 246—$25.00200KAPI only1 October 2026
36InklingThinking Machines36.5%Default#106 / 246#47$4.05524KOpen weights1 October 2026
37Gemini 3 Flash PreviewGoogle33.6%High+3 more evaluated#107 / 246—$3.001MAPI only1 October 2026
38GLM 5.2Z.ai22.8%Default#123 / 246—$3.991MOpen weights1 October 2026
39GPT-5.4 MiniOpenAI18.9%X-High+3 more evaluated#124 / 246#36$4.50400KAPI only1 October 2026
40GPT-5 ProOpenAI18.3%Default#126 / 246—$120400KAPI only1 October 2026
41GPT-5.1OpenAI17.6%High+3 more evaluated#128 / 246—$10.00400KAPI only1 October 2026
42Claude Sonnet 4.5Anthropic13.6%Thinking+4 more evaluated#132 / 246—$15.001MAPI only1 October 2026
43Kimi K2.5Moonshot AI11.8%Default#137 / 246—$2.25262KRestricted license1 October 2026
44Gemini 3.5 Flash LiteGoogle10.3%High+3 more evaluated#138 / 246#32$2.501MAPI only1 October 2026
45GPT-5OpenAI9.9%High+3 more evaluated#139 / 246—$10.00400KAPI only1 October 2026
46Claude Opus 4Anthropic8.6%Thinking+3 more evaluated#142 / 246——200KAPI only1 October 2026
47o3OpenAI6.5%High+2 more evaluated#149 / 246#57$8.00200KAPI only1 October 2026
48o4 MiniOpenAI6.1%High+2 more evaluated#151 / 246#117$4.40200KAPI only1 October 2026
49Claude Sonnet 4Anthropic5.9%Thinking+3 more evaluated#152 / 246—$15.00200KAPI only1 October 2026
50GPT-5.4 NanoOpenAI5.7%X-High+3 more evaluated#154 / 246#84$1.25400KAPI only1 October 2026
51Gemini 2.5 ProGoogle4.9%Thinking+3 more evaluated#159 / 246—$10.001MAPI only1 October 2026
52GLM 5Z.ai4.9%Default#159 / 246—$1.92205KOpen weights1 October 2026
53MiniMax M2.5MiniMax4.9%Default#159 / 246—$1.08205KRestricted license1 October 2026
54GPT-5 MiniOpenAI4.4%High+3 more evaluated#164 / 246—$2.00400KAPI only1 October 2026
55Claude Haiku 4.5Anthropic4.0%Thinking+4 more evaluated#166 / 246#75$5.00200KAPI only1 October 2026
56DeepSeek V3.2DeepSeek4.0%Default#166 / 246#66$0.42164KOpen weights1 October 2026
57o3 MiniOpenAI3.0%High+2 more evaluated#175 / 246—$4.40200KAPI only1 October 2026
58GPT-5 NanoOpenAI2.6%High+3 more evaluated#180 / 246—$0.40400KAPI only1 October 2026
59Gemini 2.5 FlashGoogle1.7%Default#193 / 246—$2.501MAPI only1 October 2026
60R1DeepSeek1.3%Default+1 more evaluated#201 / 246—$2.5064KOpen weights1 October 2026
61GPT-4.1OpenAI0.4%Default#224 / 246—$8.001MAPI only1 October 2026
62GPT-4.1 MiniOpenAI0.0%Default#228 / 246—$1.601MAPI only1 October 2026
63GPT-4.1 NanoOpenAI0.0%Default#228 / 246—$0.401MAPI only1 October 2026
64GPT-4oOpenAI0.0%Default#228 / 246—$10.00128KAPI only1 October 2026
65GPT-4o-miniOpenAI0.0%Default#228 / 246#161$0.60128KAPI only1 October 2026
66Llama 4 MaverickMeta0.0%Default#228 / 246#174$0.6521MGated access1 October 2026
67Llama 4 ScoutMeta0.0%Default#228 / 246#177$0.301.3MGated access1 October 2026

How ModelCap uses ARC-AGI-2

ARC-AGI-2 ranks models by verified ARC-AGI-2 task accuracy. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.

ARC-AGI-2 leaderboard: common questions

What is the ARC-AGI-2 benchmark?

ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. It is published by ARC Prize Foundation.

Which AI model leads ARC-AGI-2 right now?

GPT-6 Astra (OpenAI) holds the top ARC-AGI-2 score among the models ModelCap tracks, at 95.0% as of the source snapshot published 1 October 2026.

How many models are ranked on the ARC-AGI-2 leaderboard here?

67 tracked models have a published ARC-AGI-2 result on ModelCap; the source board itself lists 246 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on ARC-AGI-2?

GLM 5.3 Flash from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 65.8%.

Which model offers the best value on ARC-AGI-2?

Among the ten highest-scoring priced models, Gemini 3.8 Flash has the lowest listed output price at $3.75 per 1M tokens while scoring 89.2%.

How is ARC-AGI-2 scored?

The board ranks models by verified ARC-AGI-2 task accuracy; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does ARC-AGI-2 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the ARC-AGI-2 results?

The newest ARC-AGI-2 publication ModelCap holds is dated 1 October 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.

Explore further