Skip to content
ModelCap

Head-to-head comparison

Gemma 4 31B vs GPT-5.5

Gemma 4 31B (Google) and GPT-5.5 (OpenAI) compared on the ModelCap Index, API price, context window, provider availability, weight access and every public benchmark board they share. Figures are the same ones shown on the live rankings; nothing here is a hidden score.

Snapshot as of 22 September 2026

As of 22 September 2026, GPT-5.5 holds the stronger ModelCap Index position (#13 vs #29); Gemma 4 31B is 88.2× cheaper per output token ($0.34/1M vs $30.00/1M); GPT-5.5 offers the longer context window (1M vs 262K tokens); and Gemma 4 31B ships open weights.

Which should you choose?

Choose Gemma 4 31B if…

  • API cost matters — $0.34/1M output tokens against $30.00/1M, about 88.2× cheaper.
  • your workload is prompt-heavy — input tokens cost $0.09/1M against $5.00/1M.
  • you want to self-host: Gemma 4 31B ships open weights while GPT-5.5 is api only.
  • you want provider choice — 11 listed API providers against 3.

Choose GPT-5.5 if…

  • you want the stronger overall ModelCap Index position — #13 against #29 (82.5 vs 73.8 points).
  • you need the longer context window — 1M tokens against 262K.
  • Arena coding is your yardstick — 1520 against 1499.
  • AA τ²-Bench Telecom is your yardstick — 93.9% against 65.5%.
  • AA Humanity's Last Exam is your yardstick — 45.8% against 11.8%.
  • WildClawBench OpenClaw is your yardstick — 58.2 against 37.6.
  • you want the more recently listed model — GPT-5.5 was listed 24 April 2026, Gemma 4 31B 2 April 2026.

Gemma 4 31B vs GPT-5.5: specs, pricing and context

Specification comparison of Gemma 4 31B and GPT-5.5
FieldGemma 4 31BGoogleGPT-5.5OpenAI
ModelCap Index position#29#13
Index score73.882.5
EvidenceMeasuredMeasured
Input price / 1M tokens$0.09$5.00
Output price / 1M tokens$0.34$30.00
Context window262K tokens1M tokens
Max output tokens16K128K
API providers listed113
Weight accessOpen weightsAPI only
Input modalitiesImage, Text, VideoFile, Image, Text
First listed2 April 202624 April 2026
PublisherGoogleOpenAI

Evidence: 3 public benchmark observations across 3 boards · 7 public benchmark observations across 7 boards. Prices are the lowest listed API offer per million tokens observed on the OpenRouter catalogue.

Benchmark scores: Gemma 4 31B vs GPT-5.5

Public benchmark boards where Gemma 4 31B or GPT-5.5 has a published result
BoardGemma 4 31BGPT-5.5Leads
Arena codingArena (LMArena)14991483–151415201514–1526HighGPT-5.5
LMArena AgentLMArena5.0%4.1%–6.0%X-HighOnly one result
ARC-AGI-2ARC Prize Foundation85.0%X-HighOnly one result
ARC-AGI-3ARC Prize Foundation0.4%HighOnly one result
AA Intelligence IndexArtificial Analysis38.4intelligence-index:v4.3:xhighOnly one result
AA τ²-Bench TelecomArtificial Analysis65.5%tau2:telecom:dual-control:pass-at-1:3-repeats:source-model="Gemma 4 31B (Non-reasoning)":reasoning=false93.9%tau2:telecom:dual-control:pass-at-1:3-repeats:source-model="GPT-5.5 (xhigh)":reasoning=trueGPT-5.5
AA Humanity's Last ExamArtificial Analysis11.8%hle:may-2025:text-only-2158:no-tools:pass-at-1:source-model="Gemma 4 31B (Non-reasoning)":reasoning=false45.8%hle:may-2025:text-only-2158:no-tools:pass-at-1:source-model="GPT-5.5 (xhigh)":reasoning=trueGPT-5.5
WildClawBench OpenClawWildClawBench37.658.2GPT-5.5
Terminal-Bench 2.1 Terminus 2Terminal-Bench78.0X-HighOnly one result

Scores are the sources' own published figures for each model's best evaluated configuration; ModelCap never re-runs a benchmark.

Want a different pairing? Open the interactive comparison tool to swap either model for any current ranked language model.

Gemma 4 31B vs GPT-5.5: common questions

Is Gemma 4 31B better than GPT-5.5?

GPT-5.5 ranks higher on the ModelCap Index as of 22 September 2026: #13 against #29. That is a capability ranking built from public benchmark evidence with published uncertainty; whether it is "better" for you also depends on price, context and where you can run it. Their published uncertainty intervals overlap, so the rank difference alone does not establish a reliable capability advantage for your workload.

Is Gemma 4 31B cheaper than GPT-5.5?

Gemma 4 31B is cheaper on output tokens: $0.34/1M against $30.00/1M. Input tokens are $0.09/1M for Gemma 4 31B and $5.00/1M for GPT-5.5. Prices are the lowest listed API offer ModelCap observed, in USD per million tokens. For 1,000 requests with 2,000 input and 500 output tokens each (2M input + 0.5M output), the listed-rate estimate is $0.35 for Gemma 4 31B versus $25.00 for GPT-5.5. Gemma 4 31B costs 98.6% less in this scenario. This excludes caching, batch discounts, prompt-length tiers, tool charges and retries; verify the selected endpoint before budgeting.

Which has the bigger context window, Gemma 4 31B or GPT-5.5?

GPT-5.5 has the larger context window: 1M tokens against 262K.

Which is better for coding, Gemma 4 31B or GPT-5.5?

Arena coding: Gemma 4 31B 1499, GPT-5.5 1520 — GPT-5.5 leads.

Are Gemma 4 31B and GPT-5.5 open-weight models?

Gemma 4 31B: Open weights. GPT-5.5: API only. Open weights mean the checkpoint can be downloaded and self-hosted under its licence; API-only models are available solely through hosted endpoints.

Where can I run Gemma 4 31B and GPT-5.5?

ModelCap currently lists 11 API providers for Gemma 4 31B and 3 for GPT-5.5, from the OpenRouter catalogue snapshot the site serves; each model page lists the providers and their prices.

How current is this Gemma 4 31B vs GPT-5.5 comparison?

Every figure comes from the sealed ModelCap dataset published 22 September 2026; the page re-renders within a minute of each data refresh, and the ModelCap Index positions are the same ones shown on the live rankings.

Explore further