Skip to content
ModelCap

Benchmark leaderboard

Terminal-Bench 4.0 leaderboard

Terminal-Bench 4.0 is the official leaderboard of the terminal-work benchmark. Every published row measures a model through a vendor's own agent such as Claude Code, Codex or Grok Build, so ModelCap shows each result with its harness named and does not score the Index on it. This page lists every model ModelCap tracks with a published Terminal-Bench 4.0 result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.

Snapshot as of 11 August 2026

Tracked models with a result
0
Ranked on the ModelCap Index
0
Open-weight models listed
0

Terminal-Bench 4.0 standings (0 models)

Sorted by published score, one row per model. Source: Terminal-Bench (Laude Institute).

No tracked model currently publishes a Terminal-Bench 4.0 result in this snapshot. The standings return with the next dataset refresh that carries one; the source board itself stays available at the link above.

How ModelCap uses Terminal-Bench 4.0

Terminal-Bench 4.0 ranks models by task accuracy through the vendor's own agent. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as coding evidence alongside the other public boards; the methodology documents the weighting and the identity rules.

Terminal-Bench 4.0 leaderboard: common questions

What is the Terminal-Bench 4.0 benchmark?

Terminal-Bench 4.0 is the official leaderboard of the terminal-work benchmark. Every published row measures a model through a vendor's own agent such as Claude Code, Codex or Grok Build, so ModelCap shows each result with its harness named and does not score the Index on it. It is published by Terminal-Bench (Laude Institute).

Which AI model leads Terminal-Bench 4.0 right now?

No tracked model currently publishes a Terminal-Bench 4.0 score.

How many models are ranked on the Terminal-Bench 4.0 leaderboard here?

None in this snapshot: no tracked model has an admitted Terminal-Bench 4.0 result right now. The board repopulates from the sealed dataset as soon as one is admitted.

How is Terminal-Bench 4.0 scored?

The board ranks models by task accuracy through the vendor's own agent; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.

Does Terminal-Bench 4.0 decide the ModelCap Index rank?

Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; Terminal-Bench 4.0 contributes as coding evidence where a model has an admitted result. The methodology page documents the weighting.

How recent are the Terminal-Bench 4.0 results?

ModelCap holds no Terminal-Bench 4.0 publication in this snapshot. The page re-renders every minute from the sealed dataset and repopulates with the next admitted source row.

Explore further