Model decision surface
Compare AI models
Start with GPT-4o-mini and gpt-oss-120b, or choose any two current ranked language models. Compare capability evidence, price, context, provider availability, and weight access without pretending one field decides every use case.
Current public data
GPT-4o-mini vs gpt-oss-120b
Live dataset updated 9/22/2026, 9:31:46 PM UTC
Open 1200×630 evidence receipt| Field | GPT-4o-mini OpenAI | gpt-oss-120b OpenAI |
|---|---|---|
| ModelCap position | #139 | #131 |
| Index score | 39.7 | 43.3 |
| Evidence | Measuredpublisher-corpus-prior over 208 held-out anchors (59%); 1 specialist-board observation at 2.1% support (41%); shrunk 25.3 up toward the measured corpus | Measured3 public benchmark observations across 3 boards |
| Input / 1M | $0.15 | $0.15 |
| Output / 1M | $0.60 | $0.60 |
| Pricing status | fresh | fresh |
| Context | 128K | 131K |
| Providers | 2 | 20 |
| Weight access | API only | Open weights |
Decision facts
- GPT-4o-mini is #139; gpt-oss-120b is #131 on the same current language board.
- Index scores are 39.7 for GPT-4o-mini and 43.3 for gpt-oss-120b. Their published uncertainty intervals overlap, so the rank difference alone does not establish a reliable capability advantage for your workload.
- Both positions use Measured evidence.
- Listed output price per 1M tokens is $0.60 for GPT-4o-mini and $0.60 for gpt-oss-120b. For 1,000 requests with 2,000 input and 500 output tokens each (2M input + 0.5M output), the listed-rate estimate is $0.60 for GPT-4o-mini versus $0.60 for gpt-oss-120b. The listed-rate totals are equal. This excludes caching, batch discounts, prompt-length tiers, tool charges and retries; verify the selected endpoint before budgeting.
- Published context is 128,000 tokens for GPT-4o-mini and 131,072 for gpt-oss-120b.
- Weight access differs: GPT-4o-mini is none; gpt-oss-120b is open.
- ModelCap currently lists 2 providers for GPT-4o-mini and 20 for gpt-oss-120b.
These are separate published fields, not a synthetic winner. ModelCap does not collapse price, access, context, and capability evidence into a hidden recommendation score.
Popular comparisons
- Claude Opus 5.5 vs GPT-6 Sol#1 vs #2 on the ModelCap Index
- Claude Fable 5.1 vs Claude Opus 5.5#3 vs #1 on the ModelCap Index
- Claude Fable 5.1 vs GPT-6 Sol#3 vs #2 on the ModelCap Index
- Claude Opus 5.5 vs Grok 4.7#1 vs #4 on the ModelCap Index
- GPT-6 Sol vs Grok 4.7#2 vs #4 on the ModelCap Index
- Claude Fable 5.1 vs Grok 4.7#3 vs #4 on the ModelCap Index
- Claude Opus 5.5 vs MiMo-V2.6-Pro#1 vs #5 on the ModelCap Index
- GPT-6 Sol vs MiMo-V2.6-Pro#2 vs #5 on the ModelCap Index
- Claude Fable 5.1 vs MiMo-V2.6-Pro#3 vs #5 on the ModelCap Index
- Grok 4.7 vs MiMo-V2.6-Pro#4 vs #5 on the ModelCap Index
- Claude Opus 5.5 vs Qwen3.8 Max (0902)#1 vs #6 on the ModelCap Index
- GPT-6 Sol vs Qwen3.8 Max (0902)#2 vs #6 on the ModelCap Index
- Claude Fable 5.1 vs Qwen3.8 Max (0902)#3 vs #6 on the ModelCap Index
- Qwen3.8 Max (0902) vs Grok 4.7#6 vs #4 on the ModelCap Index
- Qwen3.8 Max (0902) vs MiMo-V2.6-Pro#6 vs #5 on the ModelCap Index
- Claude Opus 5.5 vs GPT-6 Astra#1 vs #7 on the ModelCap Index
- GPT-6 Astra vs GPT-6 Sol#7 vs #2 on the ModelCap Index
- Claude Fable 5.1 vs GPT-6 Astra#3 vs #7 on the ModelCap Index
- GPT-6 Astra vs Grok 4.7#7 vs #4 on the ModelCap Index
- GPT-6 Astra vs MiMo-V2.6-Pro#7 vs #5 on the ModelCap Index
- GPT-6 Astra vs Qwen3.8 Max (0902)#7 vs #6 on the ModelCap Index
- Claude Opus 5.5 vs Muse Spark 1.3#1 vs #8 on the ModelCap Index
- Muse Spark 1.3 vs GPT-6 Sol#8 vs #2 on the ModelCap Index
- Claude Fable 5.1 vs Muse Spark 1.3#3 vs #8 on the ModelCap Index
Each comparison page is a permanent, shareable URL with the same live figures as this tool: ModelCap Index position, API pricing, context window, provider count, weight access and every shared benchmark board.