Methodology
ModelCap is a continuously refreshed public-data index for AI models. The primary language-model rank is a capability estimate that does not require ModelCap to run or host a model. Price, popularity, downloads, likes and provider breadth never decide capability. Every input comes from a named public feed, every identity and configuration stays traceable, and modeled values are labeled separately from measurements. The public JSON and CSV snapshot, with reuse terms, is on /data.
ModelCap Rank
Current rank is a contiguous position—1, 2, 3 and so on—among current, canonical, exact language-model identities, recalculated as new evidence arrives. A usable successor can represent its product family from a defensible Modeled launch placement; older versions retain their pages and results in the catalogue. All-versions archive rank is a separate contiguous ordering across current and superseded canonical identities, rather than a frozen historical position. Distinct product tiers and sizes remain separate, and version numbers never inflate measured scores. Open-weight repositories can enter without a ModelCap-hosted evaluation, but only when their artifact identity is pinned. Image and audio catalogues stay separate because their task units are not capability-comparable. The index follows one deterministic order:
exact admitted evidence: breadth-adjusted capability point estimateverified child or quantization: explicit penalty from the exact parentno verified lineage: calibrated architecture + release-era fallbackunsupported, moving, or ambiguous identity: remain unranked
Measured rows use the existing Bayesian Evidence Score normalization. Its multi-family likelihood uses free public evaluations — 80% general preference · 12% coding · 5% agents/tools · 3% reasoning. When Overall is present it anchors the likelihood, and each observed specialist family moves that anchor by its fixed weight within a ±25-point band — an absent family contributes exactly zero, so a model is neither sunk without bound by one thin specialist board nor rewarded for never appearing on it. Overall-led models are not demoted merely because a specialist board has not tested them. The breadth adjustment applies only to specialist-only evidence. A narrow source receives a wider interval and less evidence support than broad multi-family evidence. Index v4.12 sorts capability point estimates for measured and supported modeled rows. Confidence and heuristic intervals remain separate disclosures; they are not probabilities of correctness or posterior rank probabilities. Current and All versions retain separate contiguous positions. Evidence age lowers confidence; observations below 1% combined confidence are excluded from scoring. Modeled rows show a modeled label, confidence, and interval. Static metadata alone cannot create a capability rank; unsupported or ambiguous identity also produces no rank.
The four families use Arena Overall and Artificial Analysis; controlled SWE-bench, Arena Coding and Terminal-Bench; Arena Agent, BFCL and WildClaw; and verified ARC-AGI-2 / ARC-AGI-3. Most sources use the following normalization; the versioned Artificial Analysis composite uses the calibration described below.
source score = placement and achievement, weighted per board class: 60%/40% on rating boards (Arena Overall, Arena Coding, Arena Agent) and 80%/20% on percentage boards (SWE-bench, BFCL, ARC-AGI). Placement is a confidence-interval-aware midrank among the current canonical peer field — not the raw upstream board position, which counts every configuration and superseded sibling as its own competitor. Overlapping rating intervals share a midrank, so statistical ties cannot mint order, and fields smaller than 40 models shrink toward the uninformative 50th percentile. Achievement keeps percentage boards on their published scale, while rating boards are recentred on the peer field’s own rating distribution (a logistic at the peer median, scaled by peer spread) so real rating gaps survive normalization instead of compressing into noise. A board whose newest observation across the whole corpus is older than 90 days is dormant and stops scoring for every model at once. Artificial Analysis is the exception: each reported index version is mapped from its raw units onto the Arena Overall reference scale using one shared positive Theil–Sen line. Anchors are exact admitted current models with an available provider or healthy official weights. The mapping needs at least six anchors from three publishers. Validation leaves out a whole publisher and rebuilds the reference cohort without it; the 80th percentile absolute error must be at most 16 reference-score points. Extrapolation is limited to one source interquartile range beyond the anchors, with increased error. Unsupported versions or failed calibration contribute no general-family score; their raw results remain available. Weighted held-out error widens the Index interval without lowering its point estimate. This fitted scale is not a measured Arena result or a probability of correctness.
Exact admitted independent observations replace modeled starts. Official original local releases first use attributable comparisons on their own cards when those comparisons pass support and calibration checks. Other supported estimates may use verified lineage or calibrated architecture metadata, with their disclosed penalties. Bare publisher or global corpus priors do not establish a rank under this method. Current-list membership follows usable freshly scored family successors; measured scores are never inflated to force a generation order. A later independent result can move a model down as well as up. Publisher-reported comparisons may support disclosed modeled estimates; they do not become independently measured observations. Popularity, providers, downloads and price remain separate context and cannot raise capability.
Each model has one ranked row. Arena Overall selects its highest eligible published rating across supported configurations, including Low, Medium, High, XHigh and Max. A lower effort wins when its rating is higher; Max receives no automatic preference. Specialist sources select their highest published score. Every configuration remains separate on the detail page. Versioned sources compare configurations within their latest supported board version. Identity confidence, sample size, interval width and actual source publication age determine evidence coverage.
In the active ModelCap Index, exact identity means the named catalogue product, not an assertion that every public board ran one identical endpoint configuration. The product score aggregates the admitted observations while disclosing each board’s configuration. ModelCap claims an exact endpoint configuration only after every contributing observation has a reviewed mapping; the current Index artifact does not make that claim. A Hugging Face commit is retained separately as artifact metadata and is never relabeled as the external evaluation revision.
Inside the measured-evidence subsystem, a BES result is labeled Confirmed with Arena Overall, at least 2families and evidence mass ≥ 50%; otherwise it is Provisional. Those labels describe the measured input, not a second public ordering. ModelCap keeps one capability-estimate ordering and discloses evidence support separately, with contiguous Current and All versions positions.
Filtering or sorting the table does not renumber a model. Rank stays stable within the selected scope. When scoped ranks are available, switching between Current and All versions can display a different position because those are two independently contiguous boards; retained legacy snapshots can have a single all-version scope.
The design follows published work on partial rankings with missing benchmark scores and robust cross-benchmark rank aggregation. ModelCap’s published layer is the source mapping, family weights, source-scale transforms, configuration selection and evidence adjustment—not a hidden model making subjective guesses.
Arena rating
The homepage Arena rating column shows one official external Arena result per model. When Arena publishes multiple configurations, ModelCap uses the one with the highest published Arena rating and displays that rating—not Arena’s global source rank. It does not average configurations or alter the winning result with price, context size or feature checklists. The source rank, configuration name, vote count, confidence interval and snapshot time remain on the detail page.
ModelCap joins the complete current Arena text board, then falls back to Arena’s reproducible archive if the live board is unavailable. Matches require the same publisher and full model identity. A reviewed one-way registry handles official dated names, full checkpoint names, provider renames and named configurations; fuzzy and family-name matching are never used.
A newly available model without an exact matched Arena result is labelled No matched Arena rating. It can receive a ModelCap Index position from exact verified lineage or from the active architecture-and-release-era fallback only when its support and uncertainty gates pass. Its Arena rating remains blank, and the index is explicitly labeled Inherited or Estimated rather than Measured. A named configuration such as High or Max counts as published quality evidence only for that exact configuration. When a possible cross-source identity is not reliable enough, it is labelled Identity match unresolved rather than attaching another configuration’s result.
Evaluation evidence is configuration-specific. Preview versions, reasoning levels, quantizations, and dated releases are not silently merged into a benchmark result. The catalogue folds an endpoint only when a reviewed registry declares it an identical alias of a canonical route; unreviewed lookalikes remain separate. Confidence intervals describe uncertainty in the community rating; identity-match confidence separately describes how certain the catalogue-to-evaluation name match is.
Related evaluations such as High, X-High, Max, thinking modes, or explicit token budgets appear as separate configurations on the canonical model’s detail page. The highest Arena community rating supplies the separately labelled homepage Arena value. The general-capability family uses the highest eligible Arena rating under the same model identity, with each effort's result disclosed separately. The original source result is never relabeled as a default score.
Board metrics also stay separate. For example, LMArena Agent publishes an agent score with observations and sessions—not the Arena preference rating and vote count. ModelCap labels each board and normalizes its rank before combining evidence; it never averages raw, incompatible scales.
Three evidence layers
ModelCap no longer asks one number to mean three different things. Every model page separates:
- Measured ModelCap Index
- Admitted public observations bound to one exact catalogue product, with every contributing board configuration, source field size and publication provenance disclosed separately. ModelCap did not run the model, and the Index does not claim one exact endpoint configuration across those observations.
- Inherited or Estimated ModelCap Index
- The active ModelCap Index applies an explicit configuration penalty to exact verified lineage, or uses its gated architecture-and-release-era fallback. Both remain visibly distinct from exact measured evidence, and the Index abstains when fallback support is insufficient.
- Deployment readiness
- Artifact, licence, provider and evaluation-provenance metadata. It measures how inspectable and deployable an artifact appears, never capability, safety certification or production readiness.
Hugging Face provenance
For an exact Hugging Face repository id, the sync records the immutable commit SHA, architecture, model type, total and explicitly reported active parameters, dtypes, base lineage, licence/access state, library, task, declared training datasets and repository dates. Inference-provider pricing, context, tool support, time-to-first-token and throughput stay in a separate operational feed.
Structured model-card evaluation rows are useful discovery leads, but they are not automatically trustworthy benchmark observations. Every repository row preserves the observed artifact SHA, dataset, dataset revision, configuration, task, metric, source file and verification status when published. The artifact SHA is not treated as the revision evaluated by an external board unless that source explicitly binds it. It remains quarantined and excluded from scoring until an independent admission review confirms exact identity, comparable harness and immutable dataset provenance. An unpinned dataset revision cannot pass that gate.
Cold-start posterior and validation
The cold-start estimator uses only models with admitted observed capability as anchors. Quarantined Hugging Face result values are not features. The launch proxy uses a narrow allowlist: architecture, model type, log total and active parameters, modalities, release-time proximity, exact lineage, and explicit quantization. Its calibration excludes anchors from the same publisher and lineage family so the system cannot validate a model by learning from another spelling of the same lineage.
Each result publishes its evidence state, support, score interval, expected rank, rank interval, and top-5/top-10 probabilities. Out-of-domain distance, missing static fields, weak lineage, and historical residual error widen uncertainty rather than granting an optimistic score. Unsupported or ambiguous identities still fail closed. An Estimated result remains visibly different from Measured evidence everywhere it is displayed.
Release validation is temporal: the evidence ledger freezes the first modeled posterior and compares it only with a strictly later admitted measurement. The cutover protocol additionally requires complete publisher, lineage, and architecture-family holdouts. Launch gates cover top-10 recall, frontier misses, pairwise ordering, interval calibration, group error, identity safety, and time from a complete manifest to an estimate.
This follows the practical lesson from tinyBenchmarks: a carefully calibrated subset can be informative. It also preserves the caution from research on predicting frontier model performance: extrapolation beyond observed model families is fragile and must be visible.
Launch estimate accuracy
Every modeled launch placement is frozen in the launch ledger as a prediction the moment it is published. When the model’s first independent measurement arrives, the ledger records the signed error against that frozen estimate; nothing is re-fitted after the fact. The figures below are computed from the ledger as of 22 Sept 2026, 03:10 UTC. A positive signed error means the launch estimate was optimistic.
- Prospective outcomes
- 10
- Mean absolute error
- 8.7 pts
- Signed median error
- -1.9
- 80th-percentile error
- 13.6 pts
- 80% interval coverage
- 90%
- Pairwise ordering accuracy
- 66% of 32 pairs 5+ pts apart
For context only, a retrospective replay (retrospective-2026-08-11-not-prospective) over 4 launches placed before the ledger existed shows a mean absolute error of 28.0 pts and a signed median of +22.4. Those estimates were not frozen before their outcomes were known, so they are not prospective evidence and never feed the calibration.
- · qwen/qwen3.6-27b: Launch card against resolved peers estimated 67.3, measured 23.1 on WildClawBench OpenClaw (+44.2)
- · inclusionai/ling-3.0-flash: Peer-anchored reported benchmarks estimated 60.1, measured 59.6 on Artificial Analysis Intelligence Index (+0.5)
- · z-ai/glm-5-turbo: Peer-anchored reported benchmarks estimated 80.8, measured 17.8 on WildClawBench OpenClaw (+63.0)
- · meta/muse-glimmer-30b: Peer-anchored reported benchmarks estimated 51.8, measured 55.9 on Artificial Analysis Intelligence Index, WildClawBench OpenClaw (-4.1)
Deployment readiness metadata
The readiness index is a deterministic, componentized summary of artifact reproducibility, access/licence clarity, deployability and evaluation provenance. Missing metadata lowers published coverage and stays listed on the model page. A high value does not mean the model is capable, safe, compliant, cheap or production-certified; a low value may simply mean the publisher has not exposed enough machine-readable detail.
ModelCap Core pilot
The repository now contains a versioned ModelCap Core pilot protocol and an active-sampling planner. Given a direct-run budget, the planner chooses two roles: observed residual anchors spread across the capability range, and unobserved targets prioritised by uncertainty, out-of-domain risk, board impact and lineage diversity. Unselected models receive an estimate with an interval or an abstention—they are not silently treated as tested.
The 100-item pilot allocates general, coding, agent and reasoning tasks, pins generation settings and configuration identity, repeats requests, captures request/response hashes and provider ids, and requires a sealed holdout plus inter-rater audit for subjective items. Item text remains private to reduce contamination; the public artifact is the protocol, manifest hash and aggregate evidence. Planning is local and makes no paid API calls: npm run benchmark:plan -- --budget 12.
Running the selected models is a separate measured-evidence gate because it requires provider credentials, a sealed item bank and explicit spend authority. Pilot results still cannot affect rank until the same identity, completeness, holdout and provenance checks used for external sources pass. The operator commands and private-artifact boundary are documented indocs/operations/modelcap-core-operations.md.
Market Gravity
A 0–100 composite of four broadly available signals. It measures market presence—not intelligence, and not a fictional global market cap. Market Gravity remains on model pages as transparent secondary context; it no longer determines the main ranking. See what this does not measure.
Usage
55%OpenRouter weekly popularity across the full model catalogue.
Liquidity
25%Independent providers versus the model's open or closed cohort, weighted by uptime.
Open reach
15%Hugging Face 30-day downloads for open models; neutral for closed models.
Freshness
5%Time since first public availability, on a six-month half-life.
Usage uses OpenRouter’s public weekly popularity order, which covers the full live catalogue. Because OpenRouter exposes only an ordinal without an API key, ModelCap scores that ordinal with explicit logarithmic decay and never invents a token count. Vercel’s spend, token and request exports cover only a small top-ten slice, so they are excluded from Market Gravity rather than treating missing models as zero.
Liquidity compares independent provider breadth within open and closed cohorts and modulates it with measured uptime. Open reach uses the Hugging Face 30-day download percentile for open weights; closed models receive the neutral midpoint because no comparable Hub download channel exists. Freshness has only 5% weight, enough to surface a real launch but not enough to crown one without activity.
Best value
Cheap is not the same as good. A model must first be in the top Arena preference quartile and have at least 1,000 community votes. Models without quality evidence cannot qualify.
Cost uses a published 3:1 input-to-output workload: 750,000 input tokens and 250,000 output tokens per one million combined tokens. The blended price is therefore (0.75 × input price) + (0.25 × output price).
From the eligible set, ModelCap publishes the price-performance frontier. A model is removed if another model has both an equal-or-higher quality rating and an equal-or-lower blended price, with at least one strict advantage. This prevents a poor but extremely cheap model from winning and avoids pretending that one arbitrary quality-per-dollar ratio fits every budget.
Release radar
The release radar is separate from the rankings and never changes a model's ModelCap score, position, or availability date. It scans the complete active feed under Polymarket's AI Releases tag plus a bounded scan of the broader AI feed, which catches genuine release events that lack the narrower tag. It discovers model-release families from the market questions themselves. The rolling 14-day window starts today and looks ahead, and live odds refresh every minute on Sunday and weekdays. Model names are not manually maintained.
To qualify, an item needs an active Yes/No contract with a dated “released by” question, a matching market deadline, public-release resolution rules, and a quote updated within four hours. A high-quality signal also needs a two-sided spread within five cents, at least $5,000 of liquidity, and at least $500 in trailing 24-hour volume. A current thinner contract can appear when its spread is within ten cents and liquidity is at least $100. If an otherwise current public-release contract has a one-sided order book, the radar can show its live indicative Yes price when liquidity is at least $100; it does not present that as a midpoint or use it for a date estimate. Displayed figures are market signals, not ModelCap forecasts or promised launch dates.
A market-derived date is stricter still: at least three compatible, liquid “released by” cutoffs must bracket the 25th, 50th, and 75th percentiles at useful cadence. The card then labels the interpolated 50% crossing as a market estimate and shows its 25–75% window. A lone deadline contract—or a lower-liquidity market deadline—remains a dated market projection rather than a manufactured estimate.
The homepage fills five rows with distinct qualified release families, soonest meaningful date first: a cutoff due today leads when it clears a 5% floor, then the closest date that is not a near-certain no. Odds rotate names that share a date, but do not promote a later favorite over a nearer live deadline, and do not let a 1–2% near date occupy a slot. If fewer than five families clear the floor, later dates and then sub-floor names fill the remaining rows. Related contracts such as Claude Opus, Sonnet, and Haiku are collapsed to one “Claude” family without mixing their probability ladders. If fewer than five families qualify across the source feeds, the card never invents a row. If the live market source is unavailable, the public card is withheld rather than presenting an older market view as current.
We do not turn anonymous posts, social-media speculation, or unverified leaks into public model-release claims. “Recently available” remains a separate source-backed catalogue observation.
Sources
Public, unauthenticated feeds supply catalogue, pricing, community results, venue activity and open-weight reach. The continuously refreshed validated dataset was last updated 22 Sept 2026, 03:10 UTC.
The homepage's trending-model card is a separate, same-origin live feed from Hugging Face. It shows the provider's current text-model trend order, not an invented universal activity score; the card shows the source check time and withholds retained results from the public API. Trend position and downloads never affect a model's capability rank.
- https://openrouter.ai/api/v1/modelsOpenRouterokcatalogue, pricing, context, modality, release dateFetched 22 Sept 2026, 02:49 UTC443 rows
- https://openrouter.ai/api/v1/models?sort=most-popular&output_modalities=textOpenRouter Popularityokweekly text-model popularity order (ordinal only)Fetched 22 Sept 2026, 02:49 UTC322 rows
- https://vercel.com/api/ai/leaderboard-export?dataset=models&modality=textVercel AI Gatewayokdaily top-10 request, token and spend share; latest reading and 7-day changeFetched 22 Sept 2026, 02:49 UTCPublished 22 Sept 2026, 00:00 UTC1670 rows
- https://openrouter.ai/api/v1/models/:id/endpointsOpenRouter Endpointsokper-provider pricing, measured uptime, and bounded retained provenanceFetched 22 Sept 2026, 02:49 UTC
- https://huggingface.co/api/models/:repoHugging Facepartialrevision, access and license, architecture, parameters, lineage, card metadata, reported evaluation candidates, downloads, likesFetched 22 Sept 2026, 02:49 UTC
- https://router.huggingface.co/v1/modelsHugging Face Inference Providersokexact model id, provider status, context, price, tool and structured-output support, first-token latency, throughputFetched 22 Sept 2026, 02:49 UTC134 rows
- https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetLMArenaokstyle-controlled overall community rating, confidence interval, votes, rankFetched 22 Sept 2026, 02:49 UTCPublished 13 Sept 2026, 00:00 UTCRevision d25aabda0010402 rows
- https://arena.ai/leaderboard/textArena Liveokcurrent style-controlled overall community rating, confidence interval, votes, rankFetched 22 Sept 2026, 02:49 UTCPublished 13 Sept 2026, 14:00 UTC402 rows
- https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetLMArena Agentokagent score, confidence interval, observations, sessions, rank, configurationFetched 22 Sept 2026, 02:49 UTCPublished 15 Sept 2026, 00:00 UTCRevision d25aabda001046 rows
- https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetArena Codingokstyle-controlled category rating, confidence interval, votes, rankFetched 22 Sept 2026, 02:49 UTCPublished 13 Sept 2026, 00:00 UTCRevision d25aabda0010397 rows
- https://www.swebench.com/SWE-benchokcontrolled bash-only coding score, source rank, model configurationFetched 22 Sept 2026, 02:49 UTCPublished 26 Feb 2026, 00:00 UTC47 rows
- https://www.tbench.ai/leaderboard/terminal-bench/2.1Terminal-Bench 2.1okofficial Terminus 2 task resolution, rank, effortFetched 22 Sept 2026, 02:49 UTCPublished 9 Jun 2026, 00:00 UTC5 rows
- https://www.tbench.ai/leaderboard/terminal-bench/4.0Terminal-Bench 4.0okofficial native-agent task resolution, confidence interval, trials, harness and effortFetched 22 Sept 2026, 02:49 UTCPublished 21 Sept 2026, 00:00 UTC27 rows
- https://internlm.github.io/WildClawBench/WildClawBenchokindependent OpenClaw-harness overall success rate, rankFetched 22 Sept 2026, 02:49 UTCPublished 20 Jul 2026, 00:00 UTC34 rows
- https://gorilla.cs.berkeley.edu/leaderboard.htmlBFCLokfunction and tool-use overall accuracy, rank, configurationFetched 22 Sept 2026, 02:49 UTCPublished 12 Apr 2026, 00:00 UTC109 rows
- https://arcprize.org/leaderboardARC Prize · ARC-AGI-2okverified semi-private abstract-reasoning score, source rank, reasoning effortFetched 22 Sept 2026, 02:49 UTCPublished 21 Sept 2026, 17:21 UTC214 rows
- https://arcprize.org/leaderboardARC Prize · ARC-AGI-3okverified interactive-reasoning score, source rank, reasoning effortFetched 22 Sept 2026, 02:49 UTCPublished 21 Sept 2026, 17:21 UTC39 rows
- https://epoch.ai/benchmarksEpoch AIokEpoch Capabilities Index score, evaluated configuration, confidence grade, release dateFetched 22 Sept 2026, 02:49 UTCPublished 3 Sept 2026, 00:00 UTC266 rows
- https://artificialanalysis.ai/leaderboards/modelsArtificial Analysisoknon-estimated independent composite scores and separately published HLE, Telecom and Terminal-Bench component measurements; explicit version and settingsFetched 22 Sept 2026, 02:49 UTCPublished 21 Sept 2026, 22:22 UTC649 rows
- https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/model-retirement-scheduleazureModelRetirementsokexact provider model-version retirement datesFetched 21 Sept 2026, 22:22 UTC
- https://learn.microsoft.com/en-us/azure/foundry/openai/concepts/retired-modelsazureArchivedModelRetirementsokexact provider model-version retirement datesFetched 21 Sept 2026, 22:22 UTC
- https://developers.openai.com/api/docs/deprecations.mdopenaiDirectApiRetirementsokprovider, model, retiredAt, replacementFetched 21 Sept 2026, 22:22 UTC
- https://platform.claude.com/docs/en/about-claude/model-deprecations.mdanthropicDirectApiRetirementsokprovider, model, retiredAt, replacementFetched 21 Sept 2026, 22:22 UTC
Why these benchmarks
More benchmark rows do not automatically make a better score. A source enters ModelCap only when it is public and free to read, machine-readable, dated or versioned, explicit about model configuration and evaluation harness, broad enough to compare current frontier models, and precise enough for an exact identity match. A new source also needs a declared 0–100 achievement transform before the sync will accept it.
SWE-bench Verified measures resolved real-world GitHub issues, but the result evaluates a model together with its agent scaffold. ModelCap therefore uses the official controlled bash-only board rather than mixing unrelated scaffolds. BFCL contributes structured tool-use accuracy; Arena Agent contributes independently measured behavior from real workflows; and ARC contributes only independently verified results.
Terminal-Bench 2.1, τ²-Bench Telecom and Humanity’s Last Exam are scored through the Artificial Analysis runs of each evaluation: one lab-agnostic harness measures every model within days of release, so a frontier launch is not left waiting on an upstream board. Each run feeds only the family it measures. The official Terminal-Bench 4.0 board is shown on model pages but not scored, because every published row measures a model through a vendor’s own agent. ModelCap continues to monitor LiveBench and SWE-bench Pro; they are not silently folded in when a public table mixes tool settings, scaffolds, self-reported figures or incomplete closed-model coverage.
Identity and catalogue policy
OpenRouter currently contributes 443 routing entries, while exact Hugging Face Router discovery can add independently identified repositories. After normalization ModelCap tracks 4629 model identities; 257 currently hold a public category position. These policies explain how identities are grouped, retained, ranked, or separated:
- Alias pointers
- Entries like ~anthropic/claude-opus-latest resolve to a concrete model already listed. Counting both would double every frontier release.
- Meta-routers
- openrouter/auto picks a real model per request. It is infrastructure, not a model.
- Free mirrors
- A :free entry is the same weights at a promotional price. It folds into its parent as a free-tier flag.
- Superseded versions (1468)
- A newer family member removes an older release from the contiguous Current board while preserving that exact supported identity in the separately contiguous All versions archive. Each ranked identity still needs its own evidence, verified lineage, or gated architecture-and-release-era fallback.
- Identical variants
- Only a reviewed one-way registry can declare a speed route, preset, or provider alias equivalent to one canonical product. Similar fields alone never collapse a new model.
- Adult fine-tunes
- A small denylist of roleplay-oriented publishers, so the board stays safe to open at work. Not a quality judgement.
- Non-language models
- Image, audio and video generators are ordered on separate catalogue boards from published deployment signals. Those positions are not capability ranks, and token prices do not compare across media.
- Models with no live endpoint
- The catalogue keeps listing models after every provider has dropped them. They can retain a universal Index position when exact capability support remains, while their page clearly reports that no current serving endpoint is available.
How change is measured
Every sync records an audit checkpoint of each model’s ModelCap rank, Market Gravity score and provider count, plus language-model output token price. The next compatible sync diffs those readings within the same category. Rank movement appears directly beneath the far-left rank. A methodology change starts a new movement series rather than comparing unlike scores.
A model with no prior reading is marked New and reports no change, which is the honest answer. Change pills are hidden rather than shown as zero, because a grey 0.00% would claim a measurement that was never taken.
What this does not measure
One universal definition of quality. Community preference is a useful overall signal, not proof that one model is best at coding, factual reasoning, tool use, latency-sensitive work, or a particular language. ModelCap publishes its component scores and source rows so the overall summary never hides where two models differ.
No composite proves one model is best for every task. ModelCap is an overall summary, not a claim that the top model wins every coding language, latency target, factual domain or workflow. Source-level evidence remains visible on each model page.
Throughput and latency. These are provider- and deployment-specific, not inherent model constants. ModelCap shows Hugging Face Router telemetry only for an exact repository/provider record and leaves it blank elsewhere. OpenRouter still returns those fields as null to unauthenticated clients, so the two feeds are never merged into a fictional global speed.
Release dates are catalogue dates. “First seen” is when a model appeared on a public endpoint, which is the earliest anyone could call it — usually a little after the lab’s own announcement.
Venue coverage is not the whole market. OpenRouter represents one routing venue, not direct contracts, every cloud, or consumer chat apps. Hugging Face reach and provider availability are also partial views. Market Gravity is a transparent market-presence index, not a global usage or revenue estimate.
Found something wrong? Every number on this site traces to a published source above, and the sync and scoring code is in scripts/. If a figure looks off, the arithmetic is checkable rather than a matter of trust — which is the whole point of publishing this page.