SWE-bench measures whether a model can resolve real GitHub issues by producing a patch that passes the repository's tests. The bash-only board runs every model through the same minimal agent, so the score isolates the model rather than the scaffold. This page lists every model ModelCap tracks with a published SWE-bench result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 19 February 2026
Gemini 3 Flash Preview from Google leads the SWE-bench leaderboard among the 27 tracked models with 75.8% (High configuration), per the source snapshot published 19 February 2026.
SWE-bench ranks models by resolved GitHub issues in the bash-only harness. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as coding evidence alongside the other public boards; the methodology documents the weighting and the identity rules.
SWE-bench leaderboard: common questions
What is the SWE-bench benchmark?
SWE-bench measures whether a model can resolve real GitHub issues by producing a patch that passes the repository's tests. The bash-only board runs every model through the same minimal agent, so the score isolates the model rather than the scaffold. It is published by SWE-bench (Princeton / SWE-agent team).
Which AI model leads SWE-bench right now?
Gemini 3 Flash Preview (Google) holds the top SWE-bench score among the models ModelCap tracks, at 75.8% as of the source snapshot published 19 February 2026.
How many models are ranked on the SWE-bench leaderboard here?
27 tracked models have a published SWE-bench result on ModelCap; the source board itself lists 47 entries. Each row shows the model's best evaluated configuration.
What is the best open-weight model on SWE-bench?
GLM 5 from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 72.8%.
Which model offers the best value on SWE-bench?
Among the ten highest-scoring priced models, DeepSeek V3.2 has the lowest listed output price at $0.40 per 1M tokens while scoring 70.0%.
How is SWE-bench scored?
The board ranks models by resolved GitHub issues in the bash-only harness; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does SWE-bench decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; SWE-bench contributes as coding evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the SWE-bench results?
The newest SWE-bench publication ModelCap holds is dated 19 February 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.
Which SWE-bench model ranks highest on the ModelCap Index?
DeepSeek V3.2 is the highest-scoring model on this board that currently holds a ModelCap Index position (#61).