Skip to content
ModelCap

Claude Opus 4

All-versions archive #95Arena #132API only

Anthropic·anthropic/claude-opus-4

Claude Opus 4 is ModelCap archive rank #95 (Measured), from the ModelCap Index as of 22 Sept 2026, 02:49 UTC.

Claude Opus 4 has a published context limit of 200K tokens in the ModelCap catalogue as of 22 Sept 2026, 02:49 UTC.

Claude Opus 4 weight access is api only, from ModelCap classification of published repository metadata as of 22 Sept 2026, 02:49 UTC.

Snapshot facts · download the public dataset · methodology

ModelCap Index · All-versions archive
63.8#95
Measured · 65% support · score interval 59.0–68.6
Evidence as of 21 Sept 2026, 17:21 UTC
Market Gravity
8.3
Secondary market signal

Superseded by Claude Opus 5. It no longer holds a current-board position; its all-versions archive rank is #95.

Input
Unlisted
no listed API price
Output
Unlisted
no listed API price
Context
200K
tokens
Providers
0
endpoints
Gateway spend
latest · Vercel share

Providers

Uptime measured over the last 30 minutes

No provider is currently serving this model.

Reasoning efforts

Best supported result can inform ModelCap

Published Max, X-High, High and other configurations for this model. These are sourced scores, not head-to-head essays.

Related community configurations

Human preference rating
ConfigurationRatingSource rankVotesSource
Thinking · 16K
claude-opus-4-20250514-thinking-16k
1426
CI 1422–1430
#10736KLMArena
Match 99%

Other versions (6)

ModelCap Index · All-versions archive

Methodology
63.8
Measured · 3 public benchmark observations across 3 boards
All-versions archive
#95
Public rank basis
Measured (measured)
Index support
64.6% · strong
Index interval
59.0–68.6
Rank posterior
Not available
Top 5 / Top 10 probability
Not available / Not available
Posterior as of
Snapshot timestamp unavailable
Evidence as of
21 Sept 2026, 17:21 UTC
Identity binding
Exact catalogue product · anthropic/anthropic/claude-opus-4 · aggregates disclosed benchmark configurations · endpoint configuration not claimed
Index observations used
3
Qualified Index evidence point
63.8
Bayesian Evidence Score · measured evidence input only

BES normalizes admitted public benchmark evidence for measured Index rows. Its legacy score and observables prior do not define the public language rank.

Observed capability
63.8
Evidence status
confirmed · 87% mass
General preference
63.7
Coding
70.6
Agents & tools
Reasoning
32.7
Evidence breadth
95%
Evidence coverage
65%
Market and catalogue signals · secondary, never the language Index rank
Usage (OpenRouter popularity)
Liquidity (providers × uptime)
Open reach (HF downloads)
Surface (context / tools / modalities)
88.3
Freshness
15.6
  • · Catalogue: context 200000, reasoning, tools, modalities Image/Text/File
  • · First seen 2025-05-22

These adoption and deployment observations remain context only. They do not change this model's ModelCap Index score or rank.

BES evidence selected · measured input only

The bounded 0–100 capability score blends each source’s competitive placement with its published achievement, then combines capability families. Evidence breadth adds a modest uncertainty adjustment; price and popularity are not part of benchmark-led rank.

Deployment readiness

Metadata index
30.0
limited · 37% metadata coverage
Not capability
Artifact reproducibility
0.0
Access & legal clarity
30.3
Deployability
29.0
Evaluation provenance
76.4

Missing: artifactReproducibility.revision-sha, artifactReproducibility.architecture-and-model-type, artifactReproducibility.parameter-count, artifactReproducibility.serialized-weight-format, artifactReproducibility.base-lineage, accessLegalClarity.declared-license, accessLegalClarity.resolvable-license-terms, deployability.independent-provider-records, deployability.measured-provider-uptime, deployability.published-provider-pricing, evaluationProvenance.evaluation-dataset-revisions.

A metadata completeness and deployability index—not a safety certification, quality grade, or production-readiness claim.

Market activity

Venue coverage
Market Gravity
8.3
OpenRouter weekly popularity
Not listed

Market Gravity uses OpenRouter’s full-catalogue popularity order. Sparse Vercel top-ten observations appear only when published and do not affect the score.

Arena preference

Status
Ranked
Source
LMArena
Official rank
#132 of 402
Rating
1413.9
95% confidence interval
1410–1418
Votes
43K
Source published at
13 Sept 2026, 00:00 UTC
Category
overall
Identity match confidence
99%

Market Gravity breakdown

Usage55%
0.0

OpenRouter weekly popularity across the full model catalogue.

Liquidity25%
0.0

Independent providers versus the model's open or closed cohort, weighted by uptime.

Open reach15%
50.0

Hugging Face 30-day downloads for open models; neutral for closed models.

Freshness5%
15.6

Time since first public availability, on a six-month half-life.

Specification

Context window
200K tokens
Max output
32K tokens
Inputs
Image, Text, File
Outputs
Text
Tokenizer
Claude
Cached input tokens
$1.50 / 1M
First seen on OpenRouter
22 May 2025
Knowledge cutoff
2025-01-31

Capabilities

  • Supported: Reasoning
  • Supported: Tool use
  • Not supported: Structured output
  • Not supported: Response format
  • Not supported: Moderated

Benchmark leaderboards featuring Claude Opus 4

All leaderboards
Public benchmark boards on which Claude Opus 4 has a published result
LeaderboardScoreSource rankConfigurationPublished
Arena coding leaderboardArena (LMArena)14981490–1506#70 of 397Thinking · 16K13 Sept 2026
ARC-AGI-2 leaderboardARC Prize Foundation8.6%#114 of 214Thinking21 Sept 2026

Claude Opus 4: common questions

How does Claude Opus 4 rank among AI models?

Claude Opus 4 is a superseded version and is not on the current board as of 22 September 2026. Its archive position, where one exists, is shown on the page; the newer version carries the current rank.

How much does Claude Opus 4 cost per 1M tokens?

No live API price is listed for Claude Opus 4 in the current snapshot as of 22 September 2026 because no serving endpoint is published. ModelCap shows a price only when a provider currently lists one.

What is the context window of Claude Opus 4?

Claude Opus 4 has a published context window of 200K tokens (200,000) and a maximum output of 32K tokens. Individual providers can serve less than the published maximum; the providers table lists each endpoint's own limit.

Is Claude Opus 4 open-weight?

API only. No downloadable model weights are linked in the current source data; access is through a hosted API or product.

Which benchmarks has Claude Opus 4 been evaluated on?

Claude Opus 4 has published results on 2 public boards tracked by ModelCap as of 22 September 2026: Arena coding 1498 (#70 of 397) and ARC-AGI-2 8.6% (#114 of 214). Each board page ranks every tracked model on that benchmark; the ModelCap Index combines them with published uncertainty.

When was Claude Opus 4 released?

Claude Opus 4 first appeared in the catalogue on 22 May 2025, with a published knowledge cutoff of 2025-01-31. It has since been superseded by Claude Opus 5. Rank and price movements since then are recorded on the site's changes feed.