Skip to content
ModelCap

Benchmark leaderboard

ARC-AGI-2 leaderboard

ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. This page lists every model ModelCap tracks with a published ARC-AGI-2 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 13 August 2026

GPT-5.6 Sol from OpenAI leads the ARC-AGI-2 leaderboard among the 50 tracked models with 92.5% (Max configuration), per the source snapshot published 13 August 2026.

Tracked models with a result
50
Entries on the source board
194
Open-weight models listed
4

ARC-AGI-2 standings (50 models)

Sorted by published score, one row per model. Source: ARC Prize Foundation.

ARC-AGI-2 leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1GPT-5.6 SolOpenAI92.5%Max+4 more evaluated#1 / 194#6$15.001MAPI only13 August 2026
2Claude Opus 5Anthropic90.4%Max+1 more evaluated#2 / 194#2$25.001MAPI only13 August 2026
3Claude Fable 5Anthropic89.2%Max+4 more evaluated#4 / 194#1$50.001MAPI only13 August 2026
4GPT-5.5OpenAI85.0%X-High+3 more evaluated#9 / 194#8$30.001MAPI only13 August 2026
5GPT-5.5 ProOpenAI84.6%High+1 more evaluated#10 / 194#76$1801MAPI only13 August 2026
6GPT-5.6 TerraOpenAI83.9%Max+4 more evaluated#13 / 194#12$12.001MAPI only13 August 2026
7GPT-5.4 ProOpenAI83.3%X-High#14 / 194$1801MAPI only13 August 2026
8GPT-5.4OpenAI74.0%X-High+3 more evaluated#21 / 194$15.001MAPI only13 August 2026
9Claude Opus 4.8Anthropic72.1%High+2 more evaluated#22 / 194$25.001MAPI only13 August 2026
10Gemini 3.5 FlashGoogle72.1%High+1 more evaluated#22 / 194$9.001MAPI only13 August 2026
11Claude Opus 4.6Anthropic69.2%High+3 more evaluated#26 / 194$25.001MAPI only13 August 2026
12Grok 4.6xAI67.1%X-High+3 more evaluated#31 / 194#10$6.00500KAPI only13 August 2026
13Grok 4.20xAI65.1%Reasoning#35 / 194$2.502MAPI only13 August 2026
14Claude Sonnet 4.6Anthropic60.4%High+1 more evaluated#42 / 194$15.001MAPI only13 August 2026
15Gemini 3.6 FlashGoogle60.4%High+3 more evaluated#43 / 194$3.751MAPI only13 August 2026
16Kimi K3Moonshot AI60.4%Max+2 more evaluated#43 / 194#3$15.001MRestricted license13 August 2026
17GPT-5.6 LunaOpenAI59.5%Max+4 more evaluated#46 / 194#17$1.201MAPI only13 August 2026
18GPT-5.2 ProOpenAI54.2%High+1 more evaluated#51 / 194$168400KAPI only13 August 2026
19GPT-5.2OpenAI52.9%X-High+4 more evaluated#52 / 194$14.00400KAPI only13 August 2026
20Grok 4.5xAI52.6%High+2 more evaluated#53 / 194$6.00500KAPI only13 August 2026
21Inkling SmallThinking Machines40.1%X-High+5 more evaluated#62 / 194#31$1.20524KOpen weights13 August 2026
22InklingThinking Machines36.5%Default#66 / 194#24$4.051MOpen weights13 August 2026
23Gemini 3 Flash PreviewGoogle33.6%High+3 more evaluated#67 / 194$3.001MAPI only13 August 2026
24GLM 5.2Z.ai22.8%Default#79 / 194#11$1.451MOpen weights13 August 2026
25GPT-5.4 MiniOpenAI18.9%X-High+3 more evaluated#80 / 194#23$4.50400KAPI only13 August 2026
26GPT-5 ProOpenAI18.3%Default#82 / 194$120400KAPI only13 August 2026
27GPT-5.1OpenAI17.6%High+3 more evaluated#83 / 194$10.00400KAPI only13 August 2026
28Claude Sonnet 4.5Anthropic13.6%Thinking+4 more evaluated#87 / 194$15.001MAPI only13 August 2026
29Kimi K2.5Moonshot AI11.8%Default#91 / 194$2.85262KRestricted license13 August 2026
30Gemini 3.5 Flash LiteGoogle10.3%High+3 more evaluated#92 / 194#21$2.501MAPI only13 August 2026
31GPT-5OpenAI9.9%High+3 more evaluated#93 / 194$10.00400KAPI only13 August 2026
32Claude Opus 4Anthropic8.6%Thinking+3 more evaluated#96 / 194$75.00200KAPI only13 August 2026
33o3OpenAI6.5%High+2 more evaluated#103 / 194#38$8.00200KAPI only13 August 2026
34o4 MiniOpenAI6.1%High+2 more evaluated#105 / 194#57$4.40200KAPI only13 August 2026
35Claude Sonnet 4Anthropic5.9%Thinking+3 more evaluated#106 / 194$15.001MAPI only13 August 2026
36GPT-5.4 NanoOpenAI5.7%X-High+3 more evaluated#108 / 194#46$1.25400KAPI only13 August 2026
37Gemini 2.5 ProGoogle4.9%Thinking+3 more evaluated#113 / 194$10.001MAPI only13 August 2026
38GLM 5Z.ai4.9%Default#113 / 194$1.92205KOpen weights13 August 2026
39GPT-5 MiniOpenAI4.4%High+3 more evaluated#117 / 194$2.00400KAPI only13 August 2026
40Claude Haiku 4.5Anthropic4.0%Thinking+4 more evaluated#119 / 194#45$5.00200KAPI only13 August 2026
41o3 MiniOpenAI3.0%High+2 more evaluated#128 / 194$4.40200KAPI only13 August 2026
42GPT-5 NanoOpenAI2.6%High+3 more evaluated#133 / 194$0.40400KAPI only13 August 2026
43Gemini 2.5 FlashGoogle1.7%Default#145 / 194$2.501MAPI only13 August 2026
44GPT-4.1OpenAI0.4%Default#173 / 194$8.001MAPI only13 August 2026
45GPT-4.1 MiniOpenAI0.0%Default#177 / 194$1.601MAPI only13 August 2026
46GPT-4.1 NanoOpenAI0.0%Default#177 / 194$0.401MAPI only13 August 2026
47GPT-4oOpenAI0.0%Default#177 / 194$10.00128KAPI only13 August 2026
48GPT-4o-miniOpenAI0.0%Default#177 / 194#216$0.60128KAPI only13 August 2026
49Llama 4 MaverickMeta0.0%Default#177 / 194#92$0.801MGated access13 August 2026
50Llama 4 ScoutMeta0.0%Default#177 / 194#104$0.301.3MGated access13 August 2026

How ModelCap uses ARC-AGI-2

ARC-AGI-2 ranks models by verified ARC-AGI-2 task accuracy. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as reasoning evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

ARC-AGI-2 leaderboard: common questions

What is the ARC-AGI-2 benchmark?

ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. It is published by ARC Prize Foundation.

Which AI model leads ARC-AGI-2 right now?

GPT-5.6 Sol (OpenAI) holds the top ARC-AGI-2 score among the models ModelCap tracks, at 92.5% as of the source snapshot published 13 August 2026.

How many models are ranked on the ARC-AGI-2 leaderboard here?

50 tracked models have a published ARC-AGI-2 result on ModelCap; the source board itself lists 194 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on ARC-AGI-2?

Inkling Small from Thinking Machines is the highest-scoring model with openly downloadable weights on this board, at 40.1%.

Which model offers the best value on ARC-AGI-2?

Among the ten highest-scoring priced models, Gemini 3.5 Flash has the lowest listed output price at $9.00 per 1M tokens while scoring 72.1%.

How is ARC-AGI-2 scored?

The board ranks models by verified ARC-AGI-2 task accuracy; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does ARC-AGI-2 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the ARC-AGI-2 results?

The newest ARC-AGI-2 publication ModelCap holds is dated 13 August 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Explore further