τ²-Bench Telecom measures whether an agent can resolve telecom customer-service tasks through tool calls while a simulated user talks back. Artificial Analysis runs every model through the same harness, so the pass rate compares models rather than scaffolds. This page lists every model ModelCap tracks with a published AA τ²-Bench Telecom result, its best evaluated configuration, the source's own rank, and the model's ModelCap Index position, API price and context window.
Snapshot as of 1 September 2026
GLM 5.2 from Z.ai leads the τ²-Bench Telecom (Artificial Analysis run) among the 110 tracked models with 99.1% (tau2-bench-telecom configuration), per the source snapshot published 1 September 2026.
AA τ²-Bench Telecom ranks models by τ²-Bench Telecom task pass rate. ModelCap ingests the board as published, matches each entry to a catalogue model with a reviewed identity, and shows the source's score, interval and rank unchanged. Where a model has an admitted result, it feeds the ModelCap Index as agent evidence alongside the other public boards; the methodology documents the weighting and the identity rules.
AA τ²-Bench Telecom leaderboard: common questions
What is the AA τ²-Bench Telecom benchmark?
τ²-Bench Telecom measures whether an agent can resolve telecom customer-service tasks through tool calls while a simulated user talks back. Artificial Analysis runs every model through the same harness, so the pass rate compares models rather than scaffolds. It is published by Artificial Analysis.
Which AI model leads AA τ²-Bench Telecom right now?
GLM 5.2 (Z.ai) holds the top AA τ²-Bench Telecom score among the models ModelCap tracks, at 99.1% as of the source snapshot published 1 September 2026.
How many models are ranked on the AA τ²-Bench Telecom leaderboard here?
110 tracked models have a published AA τ²-Bench Telecom result on ModelCap; the source board itself lists 439 entries. Each row shows the model's best evaluated configuration.
What is the best open-weight model on AA τ²-Bench Telecom?
GLM 5.2 from Z.ai is the highest-scoring model with openly downloadable weights on this board, at 99.1%.
Which model offers the best value on AA τ²-Bench Telecom?
Among the ten highest-scoring priced models, GLM 4.7 Flash has the lowest listed output price at $0.40 per 1M tokens while scoring 98.8%.
How is AA τ²-Bench Telecom scored?
The board ranks models by τ²-Bench Telecom task pass rate; higher scores are better. ModelCap shows the source's own score, interval and rank and never re-runs the evaluation.
Does AA τ²-Bench Telecom decide the ModelCap Index rank?
Not on its own. The ModelCap Index combines several public capability sources with published uncertainty; AA τ²-Bench Telecom contributes as agent evidence where a model has an admitted result. The methodology page documents the weighting.
How recent are the AA τ²-Bench Telecom results?
The newest AA τ²-Bench Telecom publication ModelCap holds is dated 1 September 2026. The page re-renders every minute from the sealed dataset, so it reflects the latest refresh of the source board.
Which AA τ²-Bench Telecom model ranks highest on the ModelCap Index?
Claude Fable 5 is the highest-scoring model on this board that currently holds a ModelCap Index position (#1).