ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. This page lists every model ModelCap tracks with a published ARC-AGI-2 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 13 August 2026
GPT-5.6 Sol from OpenAI leads the ARC-AGI-2 leaderboard among the 50 tracked models with 92.5% (Max configuration), per the source snapshot published 13 August 2026.
ARC-AGI-2 ranks models by verified ARC-AGI-2 task accuracy. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as reasoning evidence alongside the other public boards; the methodology documents the weighting and the identity rules.
ARC-AGI-2 leaderboard: common questions
What is the ARC-AGI-2 benchmark?
ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. It is published by ARC Prize Foundation.
Which AI model leads ARC-AGI-2 right now?
GPT-5.6 Sol (OpenAI) holds the top ARC-AGI-2 score among the models ModelCap tracks, at 92.5% as of the source snapshot published 13 August 2026.
How many models are ranked on the ARC-AGI-2 leaderboard here?
50 tracked models have a published ARC-AGI-2 result on ModelCap; the source board itself lists 194 entries. Each row shows the model's best evaluated configuration.
What is the best open-weight model on ARC-AGI-2?
Inkling Small from Thinking Machines is the highest-scoring model with openly downloadable weights on this board, at 40.1%.
Which model offers the best value on ARC-AGI-2?
Among the ten highest-scoring priced models, Gemini 3.5 Flash has the lowest listed output price at $9.00 per 1M tokens while scoring 72.1%.
How is ARC-AGI-2 scored?
The board ranks models by verified ARC-AGI-2 task accuracy; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does ARC-AGI-2 decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the ARC-AGI-2 results?
The newest ARC-AGI-2 publication ModelCap holds is dated 13 August 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.