ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. This page lists every model ModelCap tracks with a published ARC-AGI-3 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 1 October 2026
GPT-6 Astra from OpenAI leads the ARC-AGI-3 leaderboard among the 17 tracked models with 62.7% (Max configuration), per the source snapshot published 1 October 2026.
ARC-AGI-3 ranks models by verified ARC-AGI-3 task completion. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting. Read the methodology and identity rules.
ARC-AGI-3 leaderboard: common questions
What is the ARC-AGI-3 benchmark?
ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. It is published by ARC Prize Foundation.
Which AI model leads ARC-AGI-3 right now?
GPT-6 Astra (OpenAI) holds the top ARC-AGI-3 score among the models ModelCap tracks, at 62.7% as of the source snapshot published 1 October 2026.
How many models are ranked on the ARC-AGI-3 leaderboard here?
17 tracked models have a published ARC-AGI-3 result on ModelCap; the source board itself lists 79 entries. Each row shows the model's best evaluated configuration.
Which model offers the best value on ARC-AGI-3?
Among the ten highest-scoring priced models, Gemini 3.8 Flash has the lowest listed output price at $3.75 per 1M tokens while scoring 10.4%.
How is ARC-AGI-3 scored?
The board ranks models by verified ARC-AGI-3 task completion; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does ARC-AGI-3 decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the ARC-AGI-3 results?
The newest ARC-AGI-3 publication ModelCap holds is dated 1 October 2026. This page uses the published snapshot; source collection, publication and page updates happen separately.