Skip to content
ModelCap

Model decision surface

Compare AI models

Start with GPT-4o-mini and GPT-6 Luna, or choose any two current ranked language models. Compare capability evidence, price, context, provider availability, and weight access without pretending one field decides every use case.

Current public data

GPT-4o-mini vs GPT-6 Luna

Live dataset updated 9/22/2026, 10:37:09 PM UTC

Open 1200×630 evidence receipt
Factual comparison of GPT-4o-mini and GPT-6 Luna
Field
ModelCap position#139#25
Index score39.778.6
EvidenceMeasuredpublisher-corpus-prior over 209 held-out anchors (59%); 1 specialist-board observation at 2.1% support (41%); shrunk 25.3 up toward the measured corpusMeasured3 public benchmark observations across 3 boards
Input / 1M$0.15$0.10
Output / 1M$0.60$0.50
Pricing statusfreshfresh
Context128K1M
Providers22
Weight accessAPI onlyAPI only

Decision facts

  • GPT-4o-mini is #139; GPT-6 Luna is #25 on the same current language board.
  • Index scores are 39.7 for GPT-4o-mini and 78.6 for GPT-6 Luna. Their published uncertainty intervals overlap, so the rank difference alone does not establish a reliable capability advantage for your workload.
  • Both positions use Measured evidence.
  • Listed output price per 1M tokens is $0.60 for GPT-4o-mini and $0.50 for GPT-6 Luna. For 1,000 requests with 2,000 input and 500 output tokens each (2M input + 0.5M output), the listed-rate estimate is $0.60 for GPT-4o-mini versus $0.45 for GPT-6 Luna. GPT-6 Luna costs 25.0% less in this scenario. This excludes caching, batch discounts, prompt-length tiers, tool charges and retries; verify the selected endpoint before budgeting.
  • Published context is 128,000 tokens for GPT-4o-mini and 1,050,000 for GPT-6 Luna.

These are separate published fields, not a synthetic winner. ModelCap does not collapse price, access, context, and capability evidence into a hidden recommendation score.

Each comparison page is a permanent, shareable URL with the same live figures as this tool: ModelCap Index position, API pricing, context window, provider count, weight access and every shared benchmark board.