Skip to content
ModelCap

Benchmark leaderboard

ARC-AGI-3 leaderboard

ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. This page lists every model ModelCap tracks with a published ARC-AGI-3 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 13 August 2026

Claude Opus 5 from Anthropic leads the ARC-AGI-3 leaderboard among the 10 tracked models with 30.2% (High configuration), per the source snapshot published 13 August 2026.

Tracked models with a result
10
Entries on the source board
27
Open-weight models listed
0

ARC-AGI-3 standings (10 models)

Sorted by published score, one row per model. Source: ARC Prize Foundation.

ARC-AGI-3 leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Claude Opus 5Anthropic30.2%High#1 / 27#2$25.001MAPI only13 August 2026
2GPT-5.6 SolOpenAI7.8%Max+4 more evaluated#2 / 27#6$15.001MAPI only13 August 2026
3Grok 4.6xAI2.1%X-High#5 / 27#10$6.00500KAPI only13 August 2026
4Claude Opus 4.8Anthropic1.5%High#6 / 27$25.001MAPI only13 August 2026
5GPT-5.6 TerraOpenAI0.8%Max+4 more evaluated#8 / 27#12$12.001MAPI only13 August 2026
6GPT-5.5OpenAI0.4%High#12 / 27#8$30.001MAPI only13 August 2026
7Grok 4.5xAI0.3%Medium+2 more evaluated#15 / 27$6.00500KAPI only13 August 2026
8GPT-5.4OpenAI0.2%High#18 / 27$15.001MAPI only13 August 2026
9GPT-5.6 LunaOpenAI0.2%Max+4 more evaluated#19 / 27#17$1.201MAPI only13 August 2026
10Grok 4.20xAI0.1%Reasoning#24 / 27$2.502MAPI only13 August 2026

How ModelCap uses ARC-AGI-3

ARC-AGI-3 ranks models by verified ARC-AGI-3 task completion. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as reasoning evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

ARC-AGI-3 leaderboard: common questions

What is the ARC-AGI-3 benchmark?

ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. It is published by ARC Prize Foundation.

Which AI model leads ARC-AGI-3 right now?

Claude Opus 5 (Anthropic) holds the top ARC-AGI-3 score among the models ModelCap tracks, at 30.2% as of the source snapshot published 13 August 2026.

How many models are ranked on the ARC-AGI-3 leaderboard here?

10 tracked models have a published ARC-AGI-3 result on ModelCap; the source board itself lists 27 entries. Each row shows the model's best evaluated configuration.

Which model offers the best value on ARC-AGI-3?

Among the ten highest-scoring priced models, GPT-5.6 Luna has the lowest listed output price at $1.20 per 1M tokens while scoring 0.2%.

How is ARC-AGI-3 scored?

The board ranks models by verified ARC-AGI-3 task completion; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does ARC-AGI-3 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the ARC-AGI-3 results?

The newest ARC-AGI-3 publication ModelCap holds is dated 13 August 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Explore further