Every AI pricing chart shows you one number: dollars per million tokens.
That number is real. It is also close to useless for predicting your actual bill.
Two models can sit 100× apart on the price chart and land within cents of each other on your credit card — or the other way around. Here are the eight factors that decide which way it goes, with the math you can check yourself.
The price sheet at a glance
US dollars per 1M tokens, verified against official pricing pages on October 3, 2026.
| Model | Input | Output | Cache read | Notes |
|---|---|---|---|---|
| GPT-6 Luna | $0.10 | $0.50 | 10% of input | The cheap workhorse |
| DeepSeek V4.1-Flash | $0.15 | $0.60 | $0.003 | Off-peak rates; doubles in weekday peak hours (UTC) |
| Gemini 3.8 Flash | $0.75 | $3.75 | 0.1× input | Intro pricing; doubles to $1.50/$7.50 on Jan 1, 2027 |
| GPT-6.1 Sol | $2 | $10 | 10% of input | >272K-token requests re-tier the whole bill |
| Claude Sonnet 5.5 | $2 | $10 | 0.1× input | The daily-driver tier |
| Claude Opus 5.5 | $4 | $20 | 0.05× input | Step up for harder problems |
| Gemini 4 Argon | $2 | $10 | $0.10 (95% off input) | Introductory rate; Google says the regular rate will be $4/$20 (not yet in effect) |
| GPT-6 Astra | $10 | $50 | 10% of input | Frontier flagship |
| Claude Fable 5.1 | $10 | $50 | 0.025× input | Frontier flagship |
Output is ~90% of a real coding turn's cost (see Factor 2), so read the Output column first.
1. Cost per finished task, not per token
You never buy tokens. You buy done work — a fixed bug, a working feature, a test that passes.
So the only fair comparison is: how much did the finished task cost?
That math includes retries. It includes the extra turns where the model apologizes and tries again. And it includes your own time babysitting it, which is the most expensive input of all.
A worked example. Say one agent coding turn is 100K tokens of context in (cached) and 10K tokens out. At official rates:
- GPT-6 Luna: about $0.006 per turn, so $100 buys ~16,600 turns
- Claude Sonnet 5.5: about $0.12 per turn, so $100 buys ~830 turns
- GPT-6 Astra: about $0.60 per turn, so $100 buys ~166 turns
Luna buys roughly 100× more turns than Astra for the same $100. But flip it around: Astra only needs to succeed where Luna fails more than 99 times out of 100 for Astra to be the cheaper task. On hard bugs, that happens. On boilerplate, it doesn't.
What to do: for your top 3 recurring tasks, count turns per finished task per model. That one personal number beats every benchmark.
2. Input vs output pricing — coding is output-heavy
Every model charges two prices: input (what you send) and output (what it writes back). Output is almost always the expensive half.[1][2]
Look at GPT-6.1 Sol: $2 per million in, $10 per million out. Claude Sonnet 5.5 is the same $2/$10. The top tier doubles down: GPT-6 Astra and Claude Fable 5.1 charge $10 in / $50 out.
Now rerun that coding turn. 100K cached input on Sol costs $0.01. The 10K of output costs $0.10.
The output is ~90% of the turn's cost. Code generation, test writing, long explanations — it's all output. A model that looks cheap on input can be five times more expensive than you budgeted once it starts writing.
What to do: compare output prices first. Input price is the small print.
3. Verbosity: some models burn more tokens to do the same job
Price per token assumes every model spends the same number of tokens on your task. They don't.
Independent testing by Artificial Analysis found Gemini 4 Argon using roughly 62,000 output tokens per task while GPT-6 Astra used about 27,000 on the same work. Same job, more than double the tokens.[9]
Argon's rate ($2/$10) still wins that math — 62K out costs $0.62 vs Astra's $1.35 for 27K — but that is a ~2× gap, not the 5× gap the rate card suggests, and it shrank that far purely because of verbosity. Reasoning models especially will think longer, and you pay for every thought.
What to do: multiply the sticker price by the tokens-per-task, not by the token count you're imagining. Two numbers, not one.
4. Context gets re-read every turn — and long contexts get re-priced
An agent doesn't send your question once. Every turn, it re-sends the conversation: your files, the history, the tool results. A 30-turn session pays for its context roughly 30 times.
Two things keep this from bankrupting you — and one thing that can:
- Prompt caching. Cached input is 90–99% cheaper than fresh input at every major vendor. OpenAI charges 10% of input price for cached reads; Anthropic's Sonnet 5.5 cache reads cost 0.1× input; DeepSeek's off-peak cache hit is $0.003 per million against a $0.15 fresh price. If your tool doesn't cache, you are overpaying by about an order of magnitude.[1][2][4]
- The long-context cliff. OpenAI re-prices the entire request once input passes 272K tokens: 2× the input rate and 1.5× the output rate. A session that quietly grows past that line doubles its own bill mid-run, no warning shown.[1]
- Batch discounts (Anthropic's Batch API takes 50% off input and output) help only if your work can wait for async results — fine for bulk edits, useless for interactive coding.[2]
What to do: watch your context size like a fuel gauge, compact or restart sessions before they bloat, and confirm your tool actually uses caching.
5. Subscription vs API: do the break-even math, not the vibes
A subscription is a flat fee with a usage cap. The API is a meter with no cap. Neither is "cheaper" — it depends on which side of the break-even line your usage lands.
| Plan | Price | What you actually get |
|---|---|---|
| Google AI Plus | $4.99/mo | Gemini access at the budget tier (the value pick) |
| ChatGPT Plus | $20/mo | ChatGPT access; zero API credits included |
| Claude Pro | $20/mo ($17/mo billed yearly) | Claude + Claude Code for interactive work |
| Claude Max 5x / 20x | $100 / $200 per month | Same models, much higher usage caps |
| Google AI Pro | $19.99/mo | Higher Gemini limits and more storage |
| Google AI Ultra | $99.99/mo and up | Top Gemini tier |
Rough napkin math: Claude Pro is $20/month. On the API, a Sonnet 5.5 coding turn costs about $0.12. So $20 of API buys ~165 turns a month. If your agent habit is bigger than that, the subscription wins — until you hit its session caps, at which point the meter (Max tiers at $100–$200) is the only way up.
The trap is paying for both by accident: a ChatGPT Plus subscription includes zero API credits, and vice versa. Separate products, separate bills.
What to do: one flagship subscription for interactive work, API for agents and scripts, and only upgrade tiers when caps actually interrupt you — roughly once a week is the honest threshold.
6. Free tiers have a price too
Free is a real option for learning — the free tiers of ChatGPT, Claude, and the Gemini API are genuinely capable. But read what you're paying with:
- Your content. On Google's Gemini API free tier, Google says your prompts and responses may be used to improve its products. Rule: never paste secrets, client code, or anything you'd mind a stranger reading.[3]
- Your flow. Free tiers cap messages per window, then make you wait for a reset. Fine for homework, miserable mid-debugging-session.
- Your lock-in. Free tiers teach you one tool's habits. That's a cost too (see factor 8).
What to do: free for learning and throwaway projects; the moment real code or deadlines show up, pay the $20.
7. Prices have dates on them
Per-token prices are launch marketing now. They move.
The big one this year: Gemini 3.8 Flash launched at $0.75 in / $3.75 out — introductory pricing that Google states plainly on its pricing page. On January 1, 2027 it doubles to $1.50/$7.50. If Flash is your default, your Q1 bill doubles while you sleep.[3]
It's not just Google. DeepSeek halves all its rates outside its weekday UTC peak windows (01:00–04:00 and 06:00–10:00 UTC, Monday to Friday) — the same code costs half as much off-peak as it does during a weekday UTC morning peak.[4]
What to do: budget on the price you'll pay in three months, not today's headline. Check the pricing page's footnotes for the words "introductory" and "through."
8. Switching models costs money too
The cheapest model on paper can be expensive to adopt:
- Prompt rewrites. Prompts tuned for one model's habits underperform on another until you re-tune them.
- Cached prompts don't transfer. Your cache discount lives with one vendor; switching means paying full input prices while the new cache warms up.
- Tool compatibility. Agents wired to one vendor's function-calling quirks need testing, not just a new API key. (One exception: DeepSeek's API accepts both OpenAI and Anthropic request formats, so pointing existing tools at it is a config change.)
- Re-earning trust. You re-run your evals, re-learn the failure modes, re-build your instincts.
Economics favors a routing strategy instead of a single winner: cheap models (Luna at $0.10/$0.50, DeepSeek Flash) take first passes, boilerplate, and tests; a flagship handles only what the cheap model fails. You pay flagship prices on a small slice of your work instead of all of it — and if a vendor doubles prices in January, you're only re-routing one lane, not rebuilding the garage.
The 8 factors at a glance
| # | Factor | What it does to your bill | What to check |
|---|---|---|---|
| 1 | Cost per finished task | Retries and failed attempts multiply the sticker price | Turns per finished task, per model |
| 2 | Input vs output pricing | Output is ~90% of a coding turn's cost | The output column first, always |
| 3 | Verbosity | A model that writes 2× the tokens costs 2× at the same rate | Tokens per task, not per prompt |
| 4 | Context re-reads & cliffs | Uncached context re-bills every turn; >272K re-tiers the whole request | Whether your tool caches; session size |
| 5 | Subscription vs API | Flat fee vs meter — break-even depends on your real usage | Turns per month vs the sub's caps |
| 6 | Free-tier traps | Your content may train the product; caps reset mid-session | The data-use line in the fine print |
| 7 | Price-change dates | Intro prices expire; peak pricing doubles at set hours | Footnotes: "introductory," "through Dec 31" |
| 8 | Switching costs | Prompt rewrites, cold caches, tool re-testing | Routing cheap models first keeps exits cheap |
The short version
Cost per finished task beats cost per token. Output price beats input price. Tokens-per-task beats both. Cache your context, watch for the 272K cliff, match subscriptions to your real usage, treat free tiers as samples not infrastructure, budget for the price after the intro expires, and route cheap-first so switching stays cheap.
The $100 experiment says it all: the same money buys 166 turns or 16,600. The model's price tag tells you which end of that range you're shopping in — these eight factors decide where you actually land.
Disclaimers
- This post is not financial advice, and it is not an endorsement of any vendor or plan. It is a framework for comparing costs, written from the numbers available on the dates stated.
- All prices were verified against the vendors' official pricing pages on October 3, 2026. Prices, tiers, usage limits, and introductory offers change without notice — check the vendor page before you buy.
- The vendor names mentioned in this post are not sponsors of this blog, and no affiliate or referral links are included. Product names are trademarks of their owners and are used here for identification only.
- The worked examples (like the $100 coding-turns math) use the assumptions stated in the text. They are estimates to compare models, not measured benchmarks from production workloads — your own turn counts will differ, which is Factor 1 in action.
References
Pricing data courtesy of the vendors' official pricing pages (OpenAI, Anthropic, Google, DeepSeek), checked October 3, 2026; independent token-usage testing courtesy of Artificial Analysis, via The Decoder.
- OpenAI — API pricing (GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna; cached input at 10% of input; long-context rates above 272K input tokens): platform.openai.com/docs/pricing
- Anthropic — Claude model pricing (Sonnet 5.5, Opus 5.5, Fable 5.1; 0.1× cache reads standard, 0.025× on Fable 5.1, 0.05× on Opus 5.5; Batch API 50% discount): docs.anthropic.com/en/docs/about-claude/pricing
- Google — Gemini Developer API pricing (3.8 Flash $0.75/$3.75 until Dec 31, 2026, then $1.50/$7.50; cache reads at 0.1× input; free-tier content used to improve Google's products): ai.google.dev/gemini-api/docs/pricing
- DeepSeek — Models & Pricing (V4.1-Flash off-peak/peak rates, $0.003 cache-hit pricing, OpenAI/Anthropic-format base URLs): api-docs.deepseek.com/quick_start/pricing
- Google — "Introducing Gemini 4 Argon" (introductory $2/$10 with cached input 95% off; regular $4/$20 rate announced): blog.google/.../gemini-4-argon
- Anthropic — Claude plans and pricing (Pro $20/month or $17/month billed yearly, Max 5x/20x from $100): claude.com/pricing
- OpenAI Help Center — "What is ChatGPT Plus?" ($20/month; API billed separately): help.openai.com/.../what-is-chatgpt-plus
- Google AI plan tiers (AI Plus $4.99/mo, AI Pro $19.99/mo, AI Ultra from $99.99/mo): gemini.google/subscriptions
- The Decoder, reporting Artificial Analysis testing (Gemini 4 Argon ~62K vs GPT-6 Astra ~27K average output tokens per task): the-decoder.com/.../gemini-4-argon-closes-the-gap
Prices checked against official vendor pages on October 3, 2026. They will move — the factors won't.
Comments & Reactions