ModelCap
Current AI model comparison
Exact identities · public evidence
gemma 4 12B it
Position
#45
Index
61.4
Estimated
publisher-corpus-prior over 183 held-out anchors (6%); launch card against 3 resolved peers on 6 rows (94%); no cross-lab optimism probe available; shrunk 0.2 up toward the measured corpus
Comparison side 1
OpenAI
GPT-5.5 Pro
Position
#51
Index
59.3
Measured
publisher-corpus-prior over 183 held-out anchors (59%); 1 specialist-board observation at 2.1% support (41%); shrunk 14.4 toward the measured corpus
Comparison side 2
Output / 1M
Unavailable
vs $180
Context
—
vs 1M
Weights
Open weights
vs API only