ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. This page lists every model ModelCap tracks with a published ARC-AGI-2 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 1 October 2026
GPT-6 Astra from OpenAI leads the ARC-AGI-2 leaderboard among the 67 tracked models with 95.0% (Max configuration), per the source snapshot published 1 October 2026.
ARC-AGI-2 ranks models by verified ARC-AGI-2 task accuracy. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.
ARC-AGI-2 leaderboard: common questions
What is the ARC-AGI-2 benchmark?
ARC-AGI-2 measures fluid reasoning on novel visual grid puzzles that resist memorisation. The verified board reports the percentage of held-out tasks solved by each tested configuration. It is published by ARC Prize Foundation.
Which AI model leads ARC-AGI-2 right now?
GPT-6 Astra (OpenAI) holds the top ARC-AGI-2 score among the models ModelCap tracks, at 95.0% as of the source snapshot published 1 October 2026.
How many models are ranked on the ARC-AGI-2 leaderboard here?
67 tracked models have a published ARC-AGI-2 result on ModelCap; the source board itself lists 246 entries. Each row shows the model's best evaluated configuration.
What is the best open-weight model on ARC-AGI-2?
GLM 5.3 Flash from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 65.8%.
Which model offers the best value on ARC-AGI-2?
Among the ten highest-scoring priced models, Gemini 3.8 Flash has the lowest listed output price at $3.75 per 1M tokens while scoring 89.2%.
How is ARC-AGI-2 scored?
The board ranks models by verified ARC-AGI-2 task accuracy; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does ARC-AGI-2 decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-2 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the ARC-AGI-2 results?
The newest ARC-AGI-2 publication ModelCap holds is dated 1 October 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.