Choose an AI model from comparable evidence, not a single rank
A practical shortlist method: define the task, check identity and test settings, separate measurements from estimates, and record what public evidence cannot answer before choosing a route.
A leaderboard can reduce the number of models you need to consider. It cannot decide what a successful answer means for your application. A repository agent, a document extractor and an interactive writing assistant need different evidence, even when the same model appears near the top of the overall board. Start with the task, then use rank to organize a shortlist that you can explain.
Write the decision before reading the scores
Describe one unit of work and the conditions under which it is useful. For a document extractor, that might mean valid output fields, source references and a defined response when a value is absent. For a coding agent, it might mean a patch that passes the relevant checks while staying within allowed files. These are examples of acceptance rules, not claims about any model's performance.
Add the constraints that can eliminate a candidate: required input types, provider availability, usable context and output limits, latency budget, tool support and licensing. Obtain data-handling terms from the provider directly. ModelCap's catalogue and provider observations help locate the route; they do not certify it for every account or policy. A high rank does not override a failed requirement.
Use evidence close to the task
| Task | Public evidence to inspect | Still unresolved |
|---|---|---|
| Repository issue repair | A coding board with its dataset and agent setup named | Your repository, checks and allowed tools |
| Terminal automation | A terminal benchmark with its version and agent named | Your environment, permissions and recovery behavior |
| Document extraction | Relevant document evaluations and exact modality support | Your document formats and field-level acceptance rules |
| Interactive writing | Preference evidence and the observed serving route | Your audience, style and latency needs |
The coding list separates its primary multi-board tier from models with a single admitted coding result. Read that distinction before interpreting first place. The overall Index combines several kinds of public evidence; it is broader than a task-specific measurement. Popularity, downloads and a large parameter count do not substitute for a measured outcome.
Compare the result with its full label
Record the model identity, benchmark name and version, subset, metric, agent or harness, allowed tools, reasoning setting and any disclosed attempt or compute budget. A result is evidence about that combination. Two percentages on differently configured runs may not answer the same question, and an Elo rating cannot be compared numerically with a task success percentage.
The official SWE-bench site publishes distinct evaluation tracks. The Terminal-Bench leaderboard names the model and agent alongside the result. Those distinctions should survive any comparison you make. A newer benchmark version is a new measurement definition; a familiar name or reused URL does not make the old and new scores interchangeable.
Reasoning settings also need provenance. A product can have one catalogue row while its sources report several efforts or agent configurations. Read which configuration produced the admitted result; the word “Max” does not itself prove a better result or an acceptable cost. Likewise, a benchmark for an original checkpoint does not automatically measure a differently quantized serving route.
Read the evidence state and uncertainty together
ModelCap distinguishes independent measurements from modeled launch placements and publisher claims. A modeled position can be a useful early estimate, but it is not a new independent test. Inspect the model page's source links, admitted evidence, dates and method version; the evidence-layer policy explains the separation, and the row guide explains score, support and intervals.
A narrow score gap is a reason to look more closely, not proof that your application will prefer the higher row. ModelCap's intervals are heuristic uncertainty disclosures. Overlap does not establish equivalence; separation does not establish a calibrated probability that one model will win your task. Missing evidence is also not a measured failure. Write down what remains unknown instead of turning a blank cell into zero.
Keep a short decision record
- Task and acceptance: one unit of work, required outputs and failure handling.
- Candidate identity: exact model and route, with relevant settings and access constraints.
- Evidence: source URL, benchmark version, result configuration and whether it is measured or modeled.
- Cost assumptions: input and output volumes, retries and applicable provider terms.
- Unknowns and revisit trigger: unanswered application questions and the release, price or evidence change that would reopen the decision.
Use Compare for a side-by-side shortlist and the workload cost guide to calculate the token mix. If existing authorized application results are available, evaluate them against the same acceptance rule and keep failures visible. Public evidence can support the shortlist even when application evidence is unavailable; record that limitation explicitly.
Save the snapshot date and your source references with the decision. The live board can change after a release, new evidence or a method revision. A movement arrow identifies an ordering change, not its cause. Revisit when an assumption changes, rather than treating each movement as a reason to switch an application that already meets its requirements.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
- Publisher claims and independent evidence
How ModelCap distinguishes source claims, admitted measurements, and modeled estimates.
- SWE-bench official leaderboards (live)
The benchmark operator's distinct evaluation tracks and submissions. A live reference, not a historical snapshot or a guarantee about a particular application.
- Terminal-Bench leaderboard (live)
The benchmark operator's current board. Its version and harness matter when interpreting a result; this is not an archive of the older board.
Found a discrepancy? Report a correction with the article and source URL.