Why SWE-bench, BFCL and Terminal-Bench went dormant on the board, and what replaced them
Three of the coding and agent boards the Index scored have stopped publishing rows the board can use. The dormancy rule that retires them, the Terminal-Bench 4.0 change that broke the fetch, and why vendor-harness results are shown but never scored.
Independent benchmarks are the only thing that makes a ModelCap rank “measured”, and independent benchmarks have lives of their own. Maintainers move on, harnesses change, a board gets replaced by a newer version under the same URL. The Index has a rule for that, and in the summer of 2026 it fired on three boards at once. This note explains the rule, what happened to each board, and what took their place.
The dormancy rule
A board whose newest row in our corpus is more than ninety days old is dormant: it is excluded from scoring until it publishes again. Rows it already gave us stay in the dataset and keep counting for the models they measured, so nobody loses a score the day a board goes quiet; what stops is admission of new models through that board. Ninety days is long enough to ride out a maintainer's holiday and short enough that a board which has stopped for good does not keep gatekeeping a family for a year.
What happened to each board
| Board | Family | Newest row | State |
|---|---|---|---|
| SWE-bench bash-only | coding | 26 Feb 2026 | dormant since 20 May 2026; upstream stopped in February |
| BFCL v4 | agent | April 2026 | dormant since 11 July 2026 |
| Terminal-Bench 2.1 (Terminus 2) | coding | 9 Jun 2026 | retired from scoring in Score v9; served board replaced |
| WildClawBench OpenClaw | agent | 20 Jul 2026 | active; dormant on 18 Oct 2026 unless it publishes |
SWE-bench's official bash-only leaderboard simply stopped taking submissions in February; the row dates on our board show that honestly, and the benchmark page has said “February 2026” ever since. BFCL went quiet in the spring. Both are ordinary dormancy.
Terminal-Bench is the interesting one. In late August tbench.ai began serving its 4.0 board under every version URL, including the 2.1 one our fetcher read. Our parser expects a leaderboard with the pinned board's title, so it returned zero usable rows from 26 August onward and the source went to a stale-cache alert three days later. That failure was the correct behaviour: the 4.0 board is a different thing. Every row on it measures a model through a vendor's own agent, Claude Code, Codex or Grok Build, at an effort level the vendor chose. Those rows compare products, not models, and two rows for the same model can differ by ten points depending on whose harness ran it.
Why it mattered more than it looked
Coding carries 12 of the 100 points of weight and agent carries 5, so on paper three quiet boards should barely move a rank. In practice they made the specialist families hollow at the frontier: a new release could be measured on the general boards within days and then wait indefinitely for a coding or agent board that, for a closed model, might never arrive. On the day Claude Fable 5.1 launched, that was the reason it sat a few tenths of a point from its predecessor while the independent evaluations that had measured both showed a clear gap. The launch is written up in its own note.
What replaced them
- Score v9 admits three fixed-harness boards from Artificial Analysis, one per specialist family: its run of Terminal-Bench 2.1 for coding, Tau2-Bench Telecom for agents, and Humanity's Last Exam for reasoning. Artificial Analysis measures every notable model on the same harness within days of release, which is exactly the property the dormant boards had lost.
- Terminal-Bench 2.1 Terminus 2 is retired from scoring. The served Terminal-Bench 4.0 board is read from the page's embedded payload and shown as display evidence with the harness and effort named in each row's configuration, never scored. The fetch fails closed if the served board's name changes again.
- Because Artificial Analysis now carries four scored boards, v9 also caps what any single evaluator other than the Arena anchor may carry at 30% of total weight, asserted when the scoring code loads. The details are in the Score v9 note.
- Source lifecycle alerts now warn ninety and thirty days before a board's dormancy horizon and when a family's live coverage drops, and a source that has not been healthy for three days escalates to critical. A board dying quietly is what this summer taught us to watch for.
The current board list, with each board's family, weight and newest publication date, is on the methodology page, and each board has its own leaderboard page under benchmarks.