Skip to content
ModelCap

Notes

Choosing a coding model when the admitted evidence comes from one board

By ModelCapcodingguide

All 105 current coding observations at this cutoff come from Arena Coding. A refreshed shortlist separates normalized preference scores, tested efforts, overall ranks and workload prices.

The coding shortlist and the overall Index answer different questions. Coding has 12 points of family weight and its difference from general is capped at 25, so it can move an unfused general-backed Index score by at most three points either way. A model without coding evidence receives no coding adjustment. Overall rank alone therefore cannot establish coding strength.

At this cutoff, 7 of the top twenty have no coding family score: Claude Sonnet 5.5, GPT-6.1 Sol, Qwen3.8 Max (0902), LongCat 2.0, Qwen3.8 Flash, Nex N2.5 Max, Nex-N2.5-Pro. Their positions offer no measured coding result to compare.

One admitted coding board

All 105 current coding observations in the export are Arena Coding. The coding family's other registered boards, SWE-bench bash-only and Terminal-Bench 2.1 Terminus 2, have no admitted observations here. The 90-day dormancy rule removes stale boards from everyone's score together. It does not imply that all newer tracks from those organizations are stale.

The site's coding list requires at least two admitted coding boards for its primary tier. None qualifies in this cutoff; 105 models qualify for its one-board tier, whose page displays up to 100. That distinction matters: the row count in the evidence export is larger than the visible page limit. The tier sorts by the admitted coding board score, with overall rank as a tie-break.

A shortlist from the retained file

A shortlist from the retained file
Coding orderModelCoding scoreResult confidenceTested effortOverall rankBlended $/M
1Kimi K385.088.4%Max106
2GPT-6 Astra84.258.1%Max620
3Muse Spark 1.384.070.1%Max52
4Claude Opus 5.583.939.7%High28
5MiMo-V2.6-Pro83.948.1%Default80.54375
6Gemini 3.8 Flash82.985.8%High91.5
7DeepSeek V4.1 Flash82.660.9%Max120.11385
8Claude Fable 5.181.861.9%Max420
9GLM 5.3 Flash80.883.6%Default170.2375
10Gemini 3.1 Pro Preview80.4100.0%Default154.5

These are normalized 0–100 board scores, not Arena ratings or code-test pass rates. Kimi K3 leads this admitted cohort at 85.0, while GPT-6 Astra is 84.2 with lower result confidence. DeepSeek V4.1 Flash is 82.6 at $0.11385 blended, compared with Astra's $20: about 175.67 times the token price for a 1.6-point normalized difference. This does not measure cost per correct patch.

Keep source versions and intervals together

The retained coding observations name September 25 as their publication date. We separately checked the upstream live Arena coding board on September 30. Its current raw values differ from the old draft's September 13 examples. Mixing a newly read raw interval with the retained normalized score would imply a source match this analysis has not established, so those stale raw-interval comparisons are removed.

For a live decision, open the primary board and record the exact configuration, raw rating, interval, votes and publication time together. Compare candidates within that same source version. Interval overlap is useful evidence of uncertain separation; it is not a guarantee that the two systems perform identically on your code. ModelCap result confidence is an admission-and-coverage disclosure with different inputs from Arena's rating interval.

Preference and agent completion are different measurements

Arena Coding measures community preference for answers to programming prompts. It does not by itself count whether a patch passes a repository's tests. Agent benchmarks introduce another distinction: model, agent harness, tool access, effort, task version and scoring all affect a result.

Terminal-Bench 4.0 and Artificial Analysis's Terminal-Bench runs are reference evidence rather than admitted coding-family observations in this snapshot. Their usefulness depends on whether their harness resembles yours. The public evidence export used here includes scored inputs, so absence from this file does not prove a reference row is absent from a model's detail page. Consult the dated reference section there; no unverified reference score is filled in here.

Make the shortlist operational

  • Start with the coding list, then record the source's version and configuration for each finalist.
  • Compare input and output prices with the token mix your work produces, including retries and review. Use the cost guide for that calculation.
  • If you want an agent, check an evaluation of that agent and harness, rather than assuming a model-level preference result transfers unchanged.
  • Give the same small set of representative repository tasks to the finalists under a budget you explicitly choose. Record tests passed, unacceptable edits, review time and total cost.

The table makes a useful starting set. It cannot establish which model understands your repository or meets your review standard. A new independent coding board would broaden the evidence, and a future snapshot should rebuild the tiers rather than carrying these counts forward.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

  • Retained September 30 snapshot and hashes

    Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.

  • Reproduce the editorial calculations offline

    Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.

  • Published ranking method

    How the Index orders measured and modeled placements and handles succession.

  • Arena Coding primary board

    Live upstream ratings and intervals checked September 30; separate from the September 25 observations admitted to the retained snapshot.

Found a discrepancy? Report a correction with the article and source URL.

More notes