When Arena and Artificial Analysis disagree: 34 matched models
In the retained September 30 cohort, 34 models have both general observations. Their median absolute normalized gap is 5.75 points; eight differ by more than ten. Settings and calibration limit the interpretation.
Two general boards can place a model differently without either being a measurement of the other's task. Arena supplies preference evidence; Artificial Analysis supplies its versioned evaluation composite. ModelCap maps them into normalized board scores before giving Arena a 0.75 and Artificial Analysis a 0.25 share of the general family, scaled by observation confidence.
This analysis joins current-ranked models with both admitted observations in one retained file. It compares transformed scores, not raw Arena ratings against raw Intelligence Index values. All 34 matched rows, including their configuration strings, are in the downloadable calculated tables.
The observed disagreement
| Measure | Value |
|---|---|
| Matched current models | 34 |
| Median absolute gap | 5.75 points |
| Mean absolute gap | 7.43 points |
| Within 3 points | 10 / 34 |
| More than 10 points apart | 8 / 34 |
| AA higher / Arena higher / equal | 17 / 16 / 1 |
| Median signed gap (AA − Arena) | +0.15 points |
The small signed median hides substantial model-specific disagreement. It does not mean the two evaluators usually give the same score, and robust fitting does not force the arithmetic mean residual to zero.
| Model | Arena normalized | AA normalized | AA − Arena | Arena configuration | AA configuration |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 88.3 | 93.3 | +5.0 | High | intelligence-index:v4.3:reasoning-unspecified |
| GPT-6 Astra | 82 | 89.2 | +7.2 | Max | intelligence-index:v4.3:max |
| Gemini 3.8 Flash | 86.3 | 79.6 | -6.7 | High | intelligence-index:v4.3:high |
| Gemini 3.1 Pro Preview | 84.2 | 70.5 | -13.7 | Default | intelligence-index:v4.3:reasoning-unspecified |
| Grok 4.7 | 69.4 | 84.1 | +14.7 | X-High | intelligence-index:v4.3:xhigh |
| gpt-oss-20b | 32.9 | 53.6 | +20.7 | Default | intelligence-index:v4.3:high |
| Granite 4.2 8B | 25.8 | 55.3 | +29.5 | Default | intelligence-index:v4.3:reasoning-unspecified |
The largest gap is Granite 4.2 8B: 29.5 normalized points. Grok 4.7 differs by 14.7 despite naming X-High/xhigh on both observations. Gemini 3.1 Pro Preview differs by 13.7 in the opposite direction. These examples show why a single combined number can conceal evidence that matters to a particular reader.
The fitted scale is part of the comparison
The AA mapping uses a positive Theil–Sen line for each supported index version. The inspected calibration contract requires six anchors from three publishers, whole-publisher holdouts, and an 80th-percentile held-out error no larger than 16 reference points. Extrapolation is limited to one source interquartile range beyond the anchor range and increases uncertainty. A failed mapping contributes no general score.
Those rules protect against some weak mappings, but they do not turn AA into an independently measured Arena result. Because overlap models help fit the scale, a correlation calculated on the same overlap is not independent validation of agreement. We therefore report the descriptive gaps and omit the old draft's implication that a high rank correlation proves the boards agree.
Settings preserve another source of uncertainty
An Arena configuration called Default is not a claim that AA tested a default endpoint. An AA configuration containing reasoning-unspecified says the admitted provenance does not identify that setting. It cannot be renamed Default, Max or any inferred effort. Even when both sources name Max, token budgets, harnesses and source tasks can differ.
The export's configuration strings are shown as supplied. This matters for the Granite examples and gpt-oss-20b, whose admitted Arena Default and AA high are distinct settings. Different tasks and settings are plausible explanations of gaps, but these observations do not isolate a causal effect of either. Measuring that would require controlled configurations and shared tasks.
What the gap does to general score
With both observations at full confidence, general = 0.75 × Arena + 0.25 × AA. Gemini 3.1 Pro Preview's 84.2 and 70.5 become 80.775 from rounded inputs. AA lowers the general estimate by 3.425 points relative to Arena alone.
If a model starts on AA alone and later gets a full-confidence Arena result, the general-family change is 0.75 × (Arena − AA), holding the mapping and all other inputs fixed. At this cohort's median absolute gap, that is about 4.31 points in either direction. This is a sensitivity example, not a forecast interval: new votes, confidence, specialist results and refitted calibration can also move the overall Index.
Read the disagreement before using the compromise
For your finalists, inspect both source rows and their tested configurations. Decide whether preference evidence or fixed evaluation tasks better resemble your use, without treating either as a substitute for a local acceptance test. For a model on one board, the second board is new information rather than a promised increase. The score reconstruction guide shows how the admitted compromise becomes an overall score.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Retained September 30 snapshot and hashes
Exact public model and evidence exports generated September 30, 2026 at 20:54:24.517 UTC, with capture provenance and SHA-256 hashes. A fixed editorial cutoff, not live data.
- Reproduce the editorial calculations offline
Python standard-library analysis of the retained exports and minimal settled launch outcomes. No network or model inference; calculated output is retained beside the inputs.
- Published ranking method
How the Index orders measured and modeled placements and handles succession.
Found a discrepancy? Report a correction with the article and source URL.