Skip to content
ModelCap

Benchmark leaderboard

SWE-bench leaderboard

SWE-bench measures whether a model can resolve real GitHub issues by producing a patch that passes the repository's tests. The bash-only board runs every model through the same minimal agent, so the score isolates the model rather than the scaffold. This page lists every model ModelCap tracks with a published SWE-bench result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 19 February 2026

Gemini 3 Flash Preview from Google leads the SWE-bench leaderboard among the 27 tracked models with 75.8% (High configuration), per the source snapshot published 19 February 2026.

Tracked models with a result
27
Entries on the source board
47
Open-weight models listed
7

SWE-bench standings (27 models)

Sorted by published score, one row per model. Source: SWE-bench (Princeton / SWE-agent team).

SWE-bench leaderboard: every tracked model's published score, source rank, ModelCap Index position, price and context
#ModelScoreConfigurationSource rankModelCap IndexOutput $/1MContextWeightsPublished
1Gemini 3 Flash PreviewGoogle75.8%High#2 / 47$3.001MAPI only17 February 2026
2MiniMax M2.5MiniMax75.8%High#2 / 47$1.08205KRestricted license17 February 2026
3Claude Opus 4.6Anthropic75.6%Default#4 / 47$25.001MAPI only17 February 2026
4Claude Opus 4.5Anthropic74.4%Medium#5 / 47$25.00200KAPI only24 November 2025
5GLM 5Z.ai72.8%High#7 / 47$1.92205KOpen weights17 February 2026
6GPT-5.2OpenAI72.8%High#7 / 47$14.00400KAPI only17 February 2026
7GPT-5.2-CodexOpenAI72.8%Default#7 / 47$14.00400KAPI only19 February 2026
8Claude Sonnet 4.5Anthropic71.4%High+1 more evaluated#11 / 47$15.001MAPI only17 February 2026
9Kimi K2.5Moonshot AI70.8%High#12 / 47$2.25262KRestricted license17 February 2026
10DeepSeek V3.2DeepSeek70.0%High#14 / 47#61$0.40164KOpen weights17 February 2026
11Claude Haiku 4.5Anthropic66.6%High#18 / 47#70$5.00200KAPI only17 February 2026
12GPT-5.1-CodexOpenAI66.0%Medium#19 / 47$10.00400KAPI only24 November 2025
13Kimi K2 ThinkingMoonshot AI63.4%Thinking#23 / 47#47$2.50262KRestricted license10 December 2025
14MiniMax M2MiniMax61.0%Default#24 / 47$1.02205KRestricted license24 November 2025
15o3OpenAI58.4%Default#27 / 47#53$8.00200KAPI only26 July 2025
16GLM 4.6Z.ai55.4%Default#30 / 47$1.75205KOpen weights1 December 2025
17Qwen3 Coder 480B A35BQwen55.4%Default#30 / 47#104$1.00262KOpen weights2 August 2025
18GLM 4.5Z.ai54.2%Default#32 / 47$2.20131KOpen weights22 August 2025
19Devstral 2 2512Mistral AI53.8%Default#33 / 47$2.00262KRestricted license9 December 2025
20Gemini 2.5 ProGoogle53.6%Default#34 / 47$10.001MAPI only26 July 2025
21o4 MiniOpenAI45.0%Default#36 / 47#107$4.40200KAPI only26 July 2025
22GPT-4.1OpenAI39.6%Default#38 / 47$8.001MAPI only26 July 2025
23Gemini 2.5 FlashGoogle28.7%Default#40 / 47$2.501MAPI only26 July 2025
24gpt-oss-120bOpenAI26.0%Default#41 / 47#132$0.60131KOpen weights7 August 2025
25GPT-4.1 MiniOpenAI23.9%Default#42 / 47$1.601MAPI only20 July 2025
26GPT-4o (2024-11-20)OpenAI21.6%Default#43 / 47$10.00128KAPI only20 July 2025
27Qwen2.5 Coder 32B InstructQwen9.0%Default#47 / 47#209$1.0033KOpen weights3 August 2025

How ModelCap uses SWE-bench

SWE-bench ranks models by resolved GitHub issues in the bash-only harness. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as coding evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

SWE-bench leaderboard: common questions

What is the SWE-bench benchmark?

SWE-bench measures whether a model can resolve real GitHub issues by producing a patch that passes the repository's tests. The bash-only board runs every model through the same minimal agent, so the score isolates the model rather than the scaffold. It is published by SWE-bench (Princeton / SWE-agent team).

Which AI model leads SWE-bench right now?

Gemini 3 Flash Preview (Google) holds the top SWE-bench score among the models ModelCap tracks, at 75.8% as of the source snapshot published 19 February 2026.

How many models are ranked on the SWE-bench leaderboard here?

27 tracked models have a published SWE-bench result on ModelCap; the source board itself lists 47 entries. Each row shows the model's best evaluated configuration.

What is the best open-weight model on SWE-bench?

GLM 5 from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 72.8%.

Which model offers the best value on SWE-bench?

Among the ten highest-scoring priced models, DeepSeek V3.2 has the lowest listed output price at $0.40 per 1M tokens while scoring 70.0%.

How is SWE-bench scored?

The board ranks models by resolved GitHub issues in the bash-only harness; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does SWE-bench decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; SWE-bench contributes as coding evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the SWE-bench results?

The newest SWE-bench publication ModelCap holds is dated 19 February 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.

Which SWE-bench model ranks highest on the ModelCap Index?

DeepSeek V3.2 is the highest-scoring model on this board that currently holds a ModelCap Index position (#61).

Explore further