How to stop overpaying for LLM API tokens
A practical guide to what frontier model tokens actually cost, and where the cheap supply hides. Updated from live 2026 market data.
Why the sticker price is not the market price
Published prices differ by provider and processing mode. Compare current source-linked token rates below rather than assuming every aggregator charges the model vendor's list price.
Compare provider prices
USD per 1M tokens. Input and output compared separately.
Scroll sideways to see every column.
| Model | DiscountedTokens Input / Output | Compare with | Provider input / output | Input difference | Output difference |
|---|---|---|---|---|---|
| GPT-5.6 Sol gpt-5.6-sol | $0.250 / $1.500 −85% | OpenRouter ↗Checked | $2.000 / $10.000 Base token rates only; excludes cache, batch, tools, long context and platform fees. | −87% | −85% |
| GPT-5.6 Sol gpt-5.6-sol | $0.250 / $1.500 −92% | OpenAI ↗Checked | $4.000 / $20.000 Standard short-context token rates, not Batch, Flex or Fast mode. | −93% | −92% |
| GPT-5.6 Sol gpt-5.6-sol | $0.250 / $1.500 −95% | RunAPI ↗Checked | $5.000 / $30.000 Published chat token rates from the pricing page. | −95% | −95% |
| GPT-5.6 Terra gpt-5.6-terra | $0.100 / $0.600 −95% | OpenRouter ↗Checked | $2.000 / $12.000 Base token rates only; excludes cache, batch, tools, long context and platform fees. | −95% | −95% |
| GPT-5.6 Terra gpt-5.6-terra | $0.100 / $0.600 −95% | OpenAI ↗Checked | $2.000 / $12.000 Standard short-context token rates, not Batch, Flex or Fast mode. | −95% | −95% |
| GPT-5.6 Terra gpt-5.6-terra | $0.100 / $0.600 −95% | RunAPI ↗Checked | $2.000 / $12.000 Starting rates; context-dependent ranges: $2.00-$4.00 / 1M tokens input; $12.00-$18.00 / 1M tokens output. | −95% | −95% |
| GPT-5.5 gpt-5.5 | $0.250 / $1.500 −95% | OpenRouter ↗Checked | $5.000 / $30.000 Base token rates only; excludes cache, batch, tools, long context and platform fees. | −95% | −95% |
| GPT-5.5 gpt-5.5 | $0.250 / $1.500 −95% | OpenAI ↗Checked | $5.000 / $30.000 Standard short-context token rates, not Batch, Flex or Fast mode. | −95% | −95% |
| GPT-5.5 gpt-5.5 | $0.250 / $1.500 −95% | RunAPI ↗Checked | $5.000 / $30.000 Published chat token rates from the pricing page. | −95% | −95% |
| GPT-5.6 gpt-5.6 | $0.200 / $1.000 No current comparison | OpenRouter ↗Historical / unverified | $1.000 / $2.000 Base token rates only; excludes cache, batch, tools, long context and platform fees. | Not comparable | Not comparable |
| GPT-6 Astra gpt-6-astra | $0.500 / $2.500 −95% | OpenRouter ↗Checked | $10.000 / $50.000 Base token rates only; excludes cache, batch, tools, long context and platform fees. | −95% | −95% |
| GPT-6 Astra gpt-6-astra | $0.500 / $2.500 −95% | OpenAI ↗Checked | $10.000 / $50.000 Standard short-context token rates, not Batch, Flex or Fast mode. | −95% | −95% |
| GPT-6 Astra gpt-6-astra | $0.500 / $2.500 −90% | RunAPI ↗Checked | $5.000 / $25.000 Starting rates; context-dependent ranges: $5.00-$10.00 / 1M tokens input; $25.00-$37.50 / 1M tokens output. | −90% | −90% |
| Codex Auto Review codex-auto-review | $0.250 / $1.500 No current comparison | OpenRouter ↗ Not checked Historical / unverified | No verified quote / No verified quote Base token rates only; excludes cache, batch, tools, long context and platform fees. | Not comparable | Not comparable |
| Codex Auto Review codex-auto-review | $0.250 / $1.500 Higher / mixed rates | RunAPI ↗Checked | $0.200 / $1.200 Published chat token rates from the pricing page. | +25% Higher | +25% Higher |
Base token rates only; excludes cache, batch, tools, long context and platform fees.
Single discount badges use the smaller input/output saving, rounded down. Quotes older than 24 hours do not generate discount claims. Sources are checked every four hours; availability and model identity are not independently guaranteed.
Successful API requests are billed at the displayed token rates, metered to a fraction of a cent. Charges accumulate and are deducted from your prepaid balance in whole US cents, so a small request is not rounded up to a minimum charge.
We run one of those pools. Same model family, OpenAI/Anthropic/Responses-compatible endpoints, priced at the numbers below.
What you should actually pay (USD per 1M tokens)
| Model | Our input | Our output | OpenRouter list | You save |
|---|---|---|---|---|
| GPT-5.6 Sol | $0.250 | $1.500 | $2.000 / $10.000 | −85% |
| GPT-5.6 Terra | $0.100 | $0.600 | $2.000 / $12.000 | −95% |
| GPT-5.5 | $0.250 | $1.500 | $5.000 / $30.000 | −95% |
| GPT-5.6 | $0.200 | $1.000 | No verified quote / No verified quote | Reference unavailable |
| GPT-6 Astra | $0.500 | $2.500 | $10.000 / $50.000 | −95% |
| Codex Auto Review | $0.250 | $1.500 | No verified quote / No verified quote | Reference unavailable |
What about the other discount providers?
Compare source-linked RunAPI, OpenAI and OpenRouter quotes above. Published offers can conflict with a provider's pricing page; the table explains those differences. The broader Models page includes clearly labeled reference-only and historical entries, not invented discounted prices.
Three rules for cheaper inference
- Don't buy list for batch or agent workloads. Anything automated with predictable volume should never run at list price. Reseller pools are the same weights, cheaper pipe.
- Buy credits, not subscriptions. You only pay for tokens you actually consume. No unused-allowance burn.
- Keep one key for every harness. If your tooling speaks OpenAI or Anthropic, one endpoint and one key should cover Claude Code, Codex, SDKs, and agents alike.
Is a resale pool safe?
Check three things: real upstream verification (we probe model identity on every provider we resell), a stable API contract, and payment rails that clear instantly. We publish our model list and our agent-facing llms.txt so machines can audit us too.