ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. This page lists every model ModelCap tracks with a published ARC-AGI-3 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 13 August 2026
Claude Opus 5 from Anthropic leads the ARC-AGI-3 leaderboard among the 10 tracked models with 30.2% (High configuration), per the source snapshot published 13 August 2026.
ARC-AGI-3 ranks models by verified ARC-AGI-3 task completion. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as reasoning evidence alongside the other public boards; the methodology documents the weighting and the identity rules.
ARC-AGI-3 leaderboard: common questions
What is the ARC-AGI-3 benchmark?
ARC-AGI-3 moves the ARC challenge into interactive game-like environments where the model must learn the rules while acting. The verified board reports the percentage of tasks completed. It is published by ARC Prize Foundation.
Which AI model leads ARC-AGI-3 right now?
Claude Opus 5 (Anthropic) holds the top ARC-AGI-3 score among the models ModelCap tracks, at 30.2% as of the source snapshot published 13 August 2026.
How many models are ranked on the ARC-AGI-3 leaderboard here?
10 tracked models have a published ARC-AGI-3 result on ModelCap; the source board itself lists 27 entries. Each row shows the model's best evaluated configuration.
Which model offers the best value on ARC-AGI-3?
Among the ten highest-scoring priced models, GPT-5.6 Luna has the lowest listed output price at $1.20 per 1M tokens while scoring 0.2%.
How is ARC-AGI-3 scored?
The board ranks models by verified ARC-AGI-3 task completion; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does ARC-AGI-3 decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; ARC-AGI-3 contributes as reasoning evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the ARC-AGI-3 results?
The newest ARC-AGI-3 publication ModelCap holds is dated 13 August 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.