Skip to content
ModelCap

Benchmark leaderboard

ARC-AGI-3 leaderboard

ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. This page lists every model ModelCap tracks with a published ARC-AGI-3 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 1 October 2026

GPT-6 Astra from OpenAI leads the ARC-AGI-3 leaderboard among the 17 tracked models with 62.7% (Max configuration), per the source snapshot published 1 October 2026.

Tracked models with a result
17
Entries on the source board
79
Open-weight models listed
0

ARC-AGI-3 standings (17 models)

Sorted by published score, one row per model. Source: ARC Prize Foundation.

ARC-AGI-3 leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1GPT-6 AstraOpenAI62.7%Max+5 more evaluated#12 / 79#8$50.001MAPI only1 October 2026
2GPT-6.1 SolOpenAI52.7%Max+4 more evaluated#15 / 79#3$10.001MAPI only1 October 2026
3Claude Opus 5Anthropic30.2%High#20 / 79—$25.001MAPI only1 October 2026
4Gemini 3.8 FlashGoogle10.4%High+2 more evaluated#27 / 79#9$3.751MAPI only1 October 2026
5GPT-5.6 SolOpenAI7.8%Max+4 more evaluated#29 / 79—$10.001MAPI only1 October 2026
6GPT-6 SolOpenAI4.6%Max+5 more evaluated#34 / 79—$10.001MAPI only1 October 2026
7Grok 4.6xAI2.1%X-High#38 / 79—$6.00500KAPI only1 October 2026
8Claude Opus 4.8Anthropic1.5%High#41 / 79—$25.001MAPI only1 October 2026
9GPT-5.6 TerraOpenAI0.8%Max+4 more evaluated#45 / 79#22$12.001MAPI only1 October 2026
10Claude Opus 4.6Anthropic0.5%Max#48 / 79—$25.001MAPI only1 October 2026
11GPT-5.5OpenAI0.4%High#51 / 79#14$30.001MAPI only1 October 2026
12Grok 4.5xAI0.3%Medium+2 more evaluated#56 / 79—$6.00500KAPI only1 October 2026
13GPT-5.4OpenAI0.2%High#62 / 79—$15.001MAPI only1 October 2026
14GPT-6 LunaOpenAI0.2%Medium+5 more evaluated#63 / 79#31$0.501MAPI only1 October 2026
15Claude Opus 4.7Anthropic0.2%High#65 / 79—$25.001MAPI only1 October 2026
16GPT-5.6 LunaOpenAI0.2%Max+4 more evaluated#65 / 79—$1.201MAPI only1 October 2026
17Grok 4.20xAI0.1%Reasoning#73 / 79—$2.502MAPI only1 October 2026

How ModelCap uses ARC-AGI-3

ARC-AGI-3 ranks models by verified ARC-AGI-3 task completion. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.

ARC-AGI-3 leaderboard: common questions

What is the ARC-AGI-3 benchmark?

ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. It is published by ARC Prize Foundation.

Which AI model leads ARC-AGI-3 right now?

GPT-6 Astra (OpenAI) holds the top ARC-AGI-3 score among the models ModelCap tracks, at 62.7% as of the source snapshot published 1 October 2026.

How many models are ranked on the ARC-AGI-3 leaderboard here?

17 tracked models have a published ARC-AGI-3 result on ModelCap; the source board itself lists 79 entries. Each row shows the model's best evaluated configuration.

Which model offers the best value on ARC-AGI-3?

Among the ten highest-scoring priced models, Gemini 3.8 Flash has the lowest listed output price at $3.75 per 1M tokens while scoring 10.4%.

How is ARC-AGI-3 scored?

The board ranks models by verified ARC-AGI-3 task completion; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does ARC-AGI-3 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the ARC-AGI-3 results?

The newest ARC-AGI-3 publication ModelCap holds is dated 1 October 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.

Explore further