Snapshot: 2026-09-30T20:54:24.517Z Rank method: modelcap-index-v4.12-fused-launch-placement Evidence admission: modelcap-score-v10-best-supported-configuration Ranked models: 418 Measured scores from benchmark rows: 230 of 230 rebuilt within 0.15 points Family scores from benchmark rows: 562 of 562 rebuilt within 0.15 points Combined scores from their estimates: 188 of 188 rebuilt within 0.15 points Measured estimates inside combined scores: 11 of 11 rebuilt within 0.15 points Scores without rebuildable inputs: 0 Current board: 279 of 279 positions reproduced All versions: 418 of 418 positions reproduced Rank Score Rebuilt Basis Confidence Model 1 91.9 91.90 measured 20.0 Claude Sonnet 5.5 2 88.4 88.48 measured 65.4 Claude Opus 5.5 3 88.2 88.19 measured 23.0 GPT-6.1 Sol 4 86.9 86.89 measured 83.2 Claude Fable 5.1 5 84.7 84.66 measured 82.4 Muse Spark 1.3 6 83.6 83.57 measured 81.1 GPT-6 Astra 7 83.3 83.30 measured 20.0 Qwen3.8 Max (0902) 8 83.2 83.24 measured 70.8 MiMo-V2.6-Pro 9 82.7 82.67 measured 88.6 Gemini 3.8 Flash 10 82.0 81.95 measured 89.0 Kimi K3 Worked example: Claude Sonnet 5.5 general Artificial Analysis Intelligence Index: 91.9 x share 0.25 x confidence 1.000 general family: 91.90 rebuilt score: 91.90 published score: 91.9 @misc{modelcap_index, author = {{ModelCap}}, title = {Language model rankings dataset}, year = {2026}, url = {https://modelcap.ai/data}, note = {Snapshot 2026-09-30T20:54:24.517Z, rank method modelcap-index-v4.12-fused-launch-placement} }