Live per-token prices for every Claude and GPT model, side by side, plus the caching and batch discounts that decide what you really pay and what one real 20-turn conversation costs on each. Numbers pulled straight from the providers' pricing pages, checked July 2026. Then the cheapest lever most teams miss: cut your input bill up to 75% on long chats with PromptCrunch.
All prices are USD per million tokens. "Cached input" is what you pay for prompt tokens served from the provider's cache.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Anthropic (Claude) | |||
| Claude Fable 5 | $10.00 | ~$1.00 * | $50.00 |
| Claude Opus 4.8 / 4.7 / 4.6 / 4.5 | $5.00 | ~$0.50 * | $25.00 |
| Claude Opus 4.1 (legacy) | $15.00 | ~$1.50 * | $75.00 |
| Claude Sonnet 5 | $3.00 † | ~$0.30 * | $15.00 † |
| Claude Sonnet 4.6 | $3.00 | ~$0.30 * | $15.00 |
| Claude Haiku 4.5 | $1.00 | ~$0.10 * | $5.00 |
| OpenAI (GPT) | |||
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 |
| GPT-5.6 Terra | $2.50 | $0.25 | $15.00 |
| GPT-5.6 Luna | $1.00 | $0.10 | $6.00 |
| GPT-5.5 | $5.00 | $0.50 | $30.00 |
| GPT-5.5 Pro | $30.00 | Not offered | $180.00 |
| GPT-5.4 | $2.50 | $0.25 | $15.00 |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 |
| GPT-5.4 nano | $0.20 | $0.02 | $1.25 |
| GPT-5.3 Codex | $1.75 | $0.175 | $14.00 |
* Anthropic bills cache reads at roughly 0.1x the input price; the values shown are that multiple applied to the input rate. Anthropic cache writes also carry a premium (1.25x input for the 5-minute TTL, 2x for 1-hour), which OpenAI's caching does not have an equivalent of on its pricing pages.
† Sonnet 5 has introductory pricing of $2.00 input / $10.00 output through August 31, 2026.
Sources, checked July 2026: Anthropic pricing and OpenAI pricing. Prices change; confirm on the provider page before budgeting.
Rates only bite once you multiply them by real traffic. One 20-turn support conversation runs from a nickel on GPT-5.4 nano to $2.57 on Fable 5, a 50x spread for the exact same chat. Same conversation on each model (about 192,000 input tokens and 13,000 output), no caching. Only the model changes.
| Model | Cost per conversation | Per 1,000 conversations |
|---|---|---|
| Claude Fable 5 | $2.57 | $2,570 |
| GPT-5.6 Sol | $1.35 | $1,350 |
| Claude Opus 4.8 | $1.29 | $1,290 |
| Claude Sonnet 5 | $0.77 | $770 |
| GPT-5.6 Terra | $0.68 | $680 |
| GPT-5.4 | $0.68 | $680 |
| GPT-5.6 Luna | $0.27 | $270 |
| Claude Haiku 4.5 | $0.26 | $260 |
| GPT-5.4 mini | $0.20 | $200 |
| GPT-5.4 nano | $0.05 | $50 |
Assumptions: 800-token system prompt, 250-token user messages, 650-token replies, full history re-sent every turn, standard (non-intro, non-batch) rates. The Claude pricing page walks through the turn-by-turn derivation.
The sticker rates land in the same range. The discounts are where Claude and GPT split apart, and where your real bill is won or lost.
Per-token basics. Both providers bill input (what you send) and output (what the model generates) separately, per token. Output carries a large multiple: exactly 5x input across the current Claude lineup, and roughly 6x input on most current GPT models (8x on GPT-5.3 Codex). Both APIs are stateless, so a multi-turn conversation re-sends its whole history each call and pays input rates on all of it, which is why input usually dominates conversational bills despite the cheaper rate.
Caching discounts. Both providers discount repeated prompt prefixes to about 0.1x the input price, but the mechanics differ:
cache_control breakpoints, and you pay a write premium: 1.25x input for a 5-minute TTL or 2x for 1-hour. Reads then cost ~0.1x. The write premium means caching can lose money on traffic that does not re-read the prefix within the TTL.Both hinge on an exact repeated prefix. And notice what caching does not do: the history still grows every turn. Caching makes the old tokens cheaper, it never makes them disappear. That gap, the ever-growing history, is exactly what PromptCrunch closes.
Batch discounts. OpenAI documents batch processing at about half the standard rates (GPT-5.6 Sol batch input runs $2.50 versus $5.00). Anthropic offers batch too. If your work can wait for asynchronous replies, batch is the cheapest tier either provider sells.
Promotional pricing. Watch the time-boxed rates. Sonnet 5's introductory pricing (flagged in the table above) runs out August 31, 2026, then jumps to $3.00 / $15.00. Size your budget on the rate you will be paying in September, not the promo.
On a conversation that grows, most of your bill is old history you re-buy every turn. Here is how to stop paying full price for it. The three stack.
1. Compress the history before it is sent. This is the biggest recurring win, and the one the providers do not hand you. PromptCrunch shrinks the old turns of a conversation on the way to Claude or GPT, so you stop re-buying the same context on every call. On long prose chats it is worth up to 75% off your input bill. Proof and setup are in the box below.
2. Cache your stable prefix. If a big system prompt or a fixed tool block repeats on every call, provider caching discounts those repeated tokens heavily. It is built in and it stacks cleanly with compression: caching handles the parts that stay the same, compression handles the history that grows. On Anthropic, mind the write premium (1.25x for the 5-minute TTL, 2x for 1-hour); it only pays off if the prefix gets re-read inside the window.
3. Right-size the model. The conversation table above spans $0.05 to $2.57 for identical traffic, so test the cheaper models on your real workload before you reach for the flagship. This stacks too: a leaner model running a compressed prompt is the cheapest way to serve the same chat.
Caching pays off when you send the same prefix over and over. PromptCrunch pays off when the conversation grows, which is every real support thread, tutoring session, coaching chat, and companion bot, because each new turn drags the whole history along at full input price. Past roughly 20 turns our benchmarks show up to 75% fewer input tokens. On a 40-prompt Sonnet run the input bill dropped from $2.56 to $0.64, a 75% cut, for identical responses. Code, JSON, schemas, and IDs come through verbatim.
Already caching? Perfect. Requests that use provider prompt caching pass through untouched. We never break a cache hit, and we crunch the bill on everything that grows around it. Never worse than going direct, and the two savings stack.
Two lines to point any SDK at us, free to try, no card. Run the math on your own traffic in the savings estimator, or see how the two compare in compression vs caching.
$5 free credit, 100 requests a day, no card. Two lines to point any SDK at PromptCrunch, and watch the input bill fall up to 75% on your long chats. Plans are flat, starting at $29/mo.
Prices checked July 2026; always confirm on the provider's official pricing page:
Anthropic ·
OpenAI.
Deeper dives: Claude API pricing ·
OpenAI API pricing.