Compare LLM API costs for your workload, not just the cheapest token
A worked guide to input and output pricing, the 3:1 blend, workload reversals and missing charges. The examples use invented rates so the method stays useful when live prices change.
A cheap output token can be an expensive document summary. A cheap input token can be an expensive long answer. The useful comparison is the cost of the work you need, using both rates and the provider route you can actually use. ModelCap's cheapest API list is a starting point: it uses a fixed ratio to make a consistent ordering, rather than estimating every reader's bill.
Start with the units
ModelCap lists US dollars per one million tokens separately for input and output. For an ordinary text request at those rates, before other charges, the calculation is (input tokens × input rate + output tokens × output rate) / 1,000,000. Dividing each token count by one million first gives the same answer. Keep dollars per token and dollars per million tokens distinct when copying a rate from another source.
Input includes the material sent to the model, which can include instructions, retrieved documents, conversation history and tool results. Repeated steps can send that material again. Output volume also needs care: the usage record may include reasoning tokens as well as the answer you can see. OpenRouter documents native token counts and optional usage breakdowns in its API reference. Word counts and character counts are rough planning aids, not billing records.
What the 3:1 blend means
The cheapest list uses 0.75 × input rate + 0.25 × output rate. That is the cost of one million total tokens split into 750,000 input and 250,000 output tokens. It is neither the output rate nor the cost of one million input tokens plus one million output tokens. Models with a positive listed blend enter this paid-price ordering; free availability has its own list.
| Invented route | Input / 1M | Output / 1M | 3:1 blend / 1M total |
|---|---|---|---|
| A | $1.00 | $5.00 | $2.00 |
| B | $2.00 | $3.00 | $2.25 |
A sorts first by the blend even though B has the cheaper output rate. Both statements are correct; they answer different questions. The blend is deliberately a fixed comparison. A workload with a different input-to-output ratio can reverse the order.
Two workloads, two different winners
| Workload | Input / output tokens | A cost | B cost |
|---|---|---|---|
| Document summary | 12,000 / 1,000 | $0.017 | $0.027 |
| Long generated answer | 2,000 / 6,000 | $0.032 | $0.022 |
For the summary, A costs (12,000 × 1 + 1,000 × 5) / 1,000,000 = $0.017. For the long answer, A's higher output rate dominates. These two routes break even when input tokens are twice output tokens: I + 5O = 2I + 3O, so I = 2O. More input than that favors A; more output favors B. This crossover follows from the invented rates, not from a rule that all APIs follow.
Ten thousand of the summary requests would consume 120 million input and 10 million output tokens. At these unchanged example rates the token cost would be $170 for A and $270 for B. Ten thousand of the long-answer requests would instead cost $320 and $220. Estimate each workload separately before adding a monthly total; an average request can conceal an expensive minority of long documents or agent loops.
What the simple calculation leaves out
Separate uncached input, cache reads and cache writes when the route prices them differently. Check any context-length thresholds, batch or service-tier rates, and image, audio, search or tool charges that apply to the request. OpenRouter's explanation of additional charges describes why a headline rate can differ from the full charge. The displayed ModelCap token pair is not an itemized invoice and does not include every provider's billing condition.
Include billable retries and intermediate agent steps. A low rate helps only if the route produces useful work under your constraints. For existing authorized usage, total recorded cost divided by accepted completed tasks is a useful operational measure. Keep the acceptance definition fixed and report failed tasks too; dropping failures from the cost numerator makes a weak route look artificially cheap.
A comparison you can keep
- Write down the task, required output and which failures count.
- Estimate input and output volumes for each kind of request, including repeated steps.
- Record the exact model, provider route, rate units, price date and relevant billing conditions.
- Calculate a low, typical and high-volume case, with separate assumptions for caching and retries.
- Compare the estimate with existing usage receipts when available, then revisit it after a rate or workload change.
The comparison page places listed rates and capabilities beside one another, and the public dataset explains the published pricing fields and snapshot time. An unlisted price is missing information, not zero. A lowest observed offer is a catalogue observation, not a promise that your account, region or required configuration can obtain it.
Sources and further reading
ModelCap's method notes are first-party explanations, not independent benchmark measurements. Live source pages can change after publication; a current result does not establish a historical score.
- Public data and publication times
Downloadable model records, field descriptions, and snapshot freshness.
- OpenRouter usage and cost documentation
Primary documentation for native token counts, usage details and recorded generation costs; checked September 30, 2026. Provider documentation can change.
- OpenRouter explanation of additional charges
The routing service's own explanation of why a headline token price may differ from the full charge; checked September 30, 2026.
Found a discrepancy? Report a correction with the article and source URL.