ModelCap
Current AI model comparison
Exact identities · public evidence
Kwaipilot
KAT Coder V2.5 Dev
Position
#38
Index
65.9
Estimated
global-corpus-prior over 195 held-out anchors (22%); launch card against 2 resolved peers on 4 rows (78%); exceeds every named peer on 4 of 4 rows; optimism haircut 0 from cross-lab-probe; shrunk 3.7 toward the measured corpus
Comparison side 1
Poolside
Laguna S 2.1
Position
#41
Index
65.1
Estimated
5 reported rows against measured corpus ladders (72%); global-corpus-prior over 195 held-out anchors (28%); shrunk 4.8 toward the measured corpus
Comparison side 2
Output / 1M
Unavailable
vs $0.18
Context
—
vs 1M
Weights
Open weights
vs Restricted license