Skip to content
ModelCap

Notes

Choose an AI model from comparable evidence, not a single rank

By ModelCapguideevidence

A practical shortlist method: define the task, check identity and test settings, separate measurements from estimates, and record what public evidence cannot answer before choosing a route.

A leaderboard can reduce the number of models you need to consider. It cannot decide what a successful answer means for your application. A repository agent, a document extractor and an interactive writing assistant need different evidence, even when the same model appears near the top of the overall board. Start with the task, then use rank to organize a shortlist that you can explain.

Write the decision before reading the scores

Describe one unit of work and the conditions under which it is useful. For a document extractor, that might mean valid output fields, source references and a defined response when a value is absent. For a coding agent, it might mean a patch that passes the relevant checks while staying within allowed files. These are examples of acceptance rules, not claims about any model's performance.

Add the constraints that can eliminate a candidate: required input types, provider availability, usable context and output limits, latency budget, tool support and licensing. Obtain data-handling terms from the provider directly. ModelCap's catalogue and provider observations help locate the route; they do not certify it for every account or policy. A high rank does not override a failed requirement.

Use evidence close to the task

Examples of useful public evidence and the remaining application question
TaskPublic evidence to inspectStill unresolved
Repository issue repairA coding board with its dataset and agent setup namedYour repository, checks and allowed tools
Terminal automationA terminal benchmark with its version and agent namedYour environment, permissions and recovery behavior
Document extractionRelevant document evaluations and exact modality supportYour document formats and field-level acceptance rules
Interactive writingPreference evidence and the observed serving routeYour audience, style and latency needs

The coding list separates its primary multi-board tier from models with a single admitted coding result. Read that distinction before interpreting first place. The overall Index combines several kinds of public evidence; it is broader than a task-specific measurement. Popularity, downloads and a large parameter count do not substitute for a measured outcome.

Compare the result with its full label

Record the model identity, benchmark name and version, subset, metric, agent or harness, allowed tools, reasoning setting and any disclosed attempt or compute budget. A result is evidence about that combination. Two percentages on differently configured runs may not answer the same question, and an Elo rating cannot be compared numerically with a task success percentage.

The official SWE-bench site publishes distinct evaluation tracks. The Terminal-Bench leaderboard names the model and agent alongside the result. Those distinctions should survive any comparison you make. A newer benchmark version is a new measurement definition; a familiar name or reused URL does not make the old and new scores interchangeable.

Reasoning settings also need provenance. A product can have one catalogue row while its sources report several efforts or agent configurations. Read which configuration produced the admitted result; the word “Max” does not itself prove a better result or an acceptable cost. Likewise, a benchmark for an original checkpoint does not automatically measure a differently quantized serving route.

Read the evidence state and uncertainty together

ModelCap distinguishes independent measurements from modeled launch placements and publisher claims. A modeled position can be a useful early estimate, but it is not a new independent test. Inspect the model page's source links, admitted evidence, dates and method version; the evidence-layer policy explains the separation, and the row guide explains score, support and intervals.

A narrow score gap is a reason to look more closely, not proof that your application will prefer the higher row. ModelCap's intervals are heuristic uncertainty disclosures. Overlap does not establish equivalence; separation does not establish a calibrated probability that one model will win your task. Missing evidence is also not a measured failure. Write down what remains unknown instead of turning a blank cell into zero.

Keep a short decision record

  1. Task and acceptance: one unit of work, required outputs and failure handling.
  2. Candidate identity: exact model and route, with relevant settings and access constraints.
  3. Evidence: source URL, benchmark version, result configuration and whether it is measured or modeled.
  4. Cost assumptions: input and output volumes, retries and applicable provider terms.
  5. Unknowns and revisit trigger: unanswered application questions and the release, price or evidence change that would reopen the decision.

Use Compare for a side-by-side shortlist and the workload cost guide to calculate the token mix. If existing authorized application results are available, evaluate them against the same acceptance rule and keep failures visible. Public evidence can support the shortlist even when application evidence is unavailable; record that limitation explicitly.

Save the snapshot date and your source references with the decision. The live board can change after a release, new evidence or a method revision. A movement arrow identifies an ordering change, not its cause. Revisit when an assumption changes, rather than treating each movement as a reason to switch an application that already meets its requirements.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

Found a discrepancy? Report a correction with the article and source URL.

More notes