Skip to content
ModelCap

Notes

When Arena and Artificial Analysis disagree: 34 matched models

By ModelCapbenchmarksmethodology

In the retained September 30 cohort, 34 models have both general observations. Their median absolute normalized gap is 5.75 points; eight differ by more than ten. Settings and calibration limit the interpretation.

Two general boards can place a model differently without either being a measurement of the other's task. Arena supplies preference evidence; Artificial Analysis supplies its versioned evaluation composite. ModelCap maps them into normalized board scores before giving Arena a 0.75 and Artificial Analysis a 0.25 share of the general family, scaled by observation confidence.

This analysis joins current-ranked models with both admitted observations in one retained file. It compares transformed scores, not raw Arena ratings against raw Intelligence Index values. All 34 matched rows, including their configuration strings, are in the downloadable calculated tables.

The observed disagreement

The observed disagreement
MeasureValue
Matched current models34
Median absolute gap5.75 points
Mean absolute gap7.43 points
Within 3 points10 / 34
More than 10 points apart8 / 34
AA higher / Arena higher / equal17 / 16 / 1
Median signed gap (AA − Arena)+0.15 points

The small signed median hides substantial model-specific disagreement. It does not mean the two evaluators usually give the same score, and robust fitting does not force the arithmetic mean residual to zero.

The observed disagreement
ModelArena normalizedAA normalizedAA − ArenaArena configurationAA configuration
Claude Opus 5.588.393.3+5.0Highintelligence-index:v4.3:reasoning-unspecified
GPT-6 Astra8289.2+7.2Maxintelligence-index:v4.3:max
Gemini 3.8 Flash86.379.6-6.7Highintelligence-index:v4.3:high
Gemini 3.1 Pro Preview84.270.5-13.7Defaultintelligence-index:v4.3:reasoning-unspecified
Grok 4.769.484.1+14.7X-Highintelligence-index:v4.3:xhigh
gpt-oss-20b32.953.6+20.7Defaultintelligence-index:v4.3:high
Granite 4.2 8B25.855.3+29.5Defaultintelligence-index:v4.3:reasoning-unspecified

The largest gap is Granite 4.2 8B: 29.5 normalized points. Grok 4.7 differs by 14.7 despite naming X-High/xhigh on both observations. Gemini 3.1 Pro Preview differs by 13.7 in the opposite direction. These examples show why a single combined number can conceal evidence that matters to a particular reader.

The fitted scale is part of the comparison

The AA mapping uses a positive Theil–Sen line for each supported index version. The inspected calibration contract requires six anchors from three publishers, whole-publisher holdouts, and an 80th-percentile held-out error no larger than 16 reference points. Extrapolation is limited to one source interquartile range beyond the anchor range and increases uncertainty. A failed mapping contributes no general score.

Those rules protect against some weak mappings, but they do not turn AA into an independently measured Arena result. Because overlap models help fit the scale, a correlation calculated on the same overlap is not independent validation of agreement. We therefore report the descriptive gaps and omit the old draft's implication that a high rank correlation proves the boards agree.

Settings preserve another source of uncertainty

An Arena configuration called Default is not a claim that AA tested a default endpoint. An AA configuration containing reasoning-unspecified says the admitted provenance does not identify that setting. It cannot be renamed Default, Max or any inferred effort. Even when both sources name Max, token budgets, harnesses and source tasks can differ.

The export's configuration strings are shown as supplied. This matters for the Granite examples and gpt-oss-20b, whose admitted Arena Default and AA high are distinct settings. Different tasks and settings are plausible explanations of gaps, but these observations do not isolate a causal effect of either. Measuring that would require controlled configurations and shared tasks.

What the gap does to general score

With both observations at full confidence, general = 0.75 × Arena + 0.25 × AA. Gemini 3.1 Pro Preview's 84.2 and 70.5 become 80.775 from rounded inputs. AA lowers the general estimate by 3.425 points relative to Arena alone.

If a model starts on AA alone and later gets a full-confidence Arena result, the general-family change is 0.75 × (Arena − AA), holding the mapping and all other inputs fixed. At this cohort's median absolute gap, that is about 4.31 points in either direction. This is a sensitivity example, not a forecast interval: new votes, confidence, specialist results and refitted calibration can also move the overall Index.

Read the disagreement before using the compromise

For your finalists, inspect both source rows and their tested configurations. Decide whether preference evidence or fixed evaluation tasks better resemble your use, without treating either as a substitute for a local acceptance test. For a model on one board, the second board is new information rather than a promised increase. The score reconstruction guide shows how the admitted compromise becomes an overall score.

Sources and further reading

ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.

Found a discrepancy? Report a correction with the article and source URL.

More notes