Anthropic shipped Claude Fable 5.1 on 1 September 2026. Within a day the usual comparison posts appeared, all benchmarks and no invoices. So here is the version that matters when you are the one paying: what the three flagship models actually cost per million tokens right now, and which one you should put your money on for which job.
This is a spending decision, not a leaderboard. Below are the real published rates, the trap that doubles one of them without warning, and a decision rule you can apply this week.
What the three flagship models cost per million tokens
As of September 2026, Claude Fable 5.1 is US$10 per million input tokens and US$50 per million output tokens. GPT-5.5 is US$5 input and US$30 output. Gemini 3.1 Pro is US$2 input and US$12 output for standard prompts. That is a five-fold input spread and a four-fold output spread across the same tier of model.
Two numbers sit outside that table and change decisions more than the headline rates do.
Claude Fable 5.1 cache reads are US$0.25 per million tokens. That is 75% less than Fable 5 charged, and it is the reason Anthropic describes typical workloads as roughly 25% cheaper and highly agentic workloads as up to about 45% cheaper. If your workload re-reads the same long document, style guide or codebase on every turn, the cache rate is your real rate, not the input rate.
GPT-5.5 Pro is US$30 input and US$180 output. This is a separate research-grade tier, not the default. It is six times the input cost and six times the output cost of standard GPT-5.5. If you have ever been surprised by an OpenAI bill, check which of these two you were actually calling.
The 200,000-token cliff that doubles your Gemini bill
Gemini 3.1 Pro's US$2 input rate applies only below 200,000 tokens of context in a single prompt. Cross that line and input roughly doubles to US$4 per million while output climbs from US$12 to US$18. The cheapest model on the list stops being the cheapest exactly when you start using the long context you chose it for.
This catches people because the long context window is Gemini's headline feature. The pitch is a million tokens; the price break sits at a fifth of that. If your workflow is "dump the entire quarterly report in and ask questions", you are on the far side of the cliff and paying US$4 and US$18, which narrows the gap against GPT-5.5 considerably.
The practical response is not to avoid long context. It is to know which of your prompts cross 200,000 tokens and split the ones that do not need to. Chunking a 300,000-token document into two 150,000-token passes keeps you on the cheap side of the line, at the cost of losing cross-document reasoning. Whether that trade is worth it depends entirely on the task, which is the honest answer nobody wants.
What you actually pay as an individual, not an API caller
If you are a marketer, writer or operations manager rather than a developer, per-token rates are background information. The consumer tiers are what hit your card, and they have converged on roughly US$20 a month each: ChatGPT Plus at US$20, Claude Pro at US$20 and Gemini's paid tier at US$19.99. In Hong Kong terms that is about HK$155 a month per subscription.
The convergence is the important fact. Because the subscription prices are effectively identical, price is not a tie-breaker at the individual level, capability-per-task is. This is the opposite of the API picture, where a five-fold cost spread makes price the dominant variable. Practitioners keep importing API-level cost anxiety into a subscription decision where it does not apply.
Two subscriptions cost about HK$310 a month. That is a real number for a freelancer and a rounding error for a company. If you are billing client work through these tools, running two is usually the better decision, and we will come back to which two.
The decision rule: match the model to the task, not the benchmark
Ignore the composite leaderboard scores. They average across task types you do not perform. The useful rule is to sort your own work into three buckets and assign a model to each: long-form written output, high-volume cheap throughput, and large-context analysis.
Long-form written output. Claude has the strongest reputation for natural long-form prose, and at US$50 per million output tokens it is also the most expensive place to generate a lot of words. That combination points to a specific use: final drafts and high-stakes writing, not first drafts and bulk content. Use the expensive model where the output is the deliverable.
High-volume cheap throughput. Classification, tagging, summarising a hundred support tickets, drafting variants. Gemini 3.1 Pro's US$2 and US$12 rates win this outright as long as each prompt stays under the 200,000-token line, which for this kind of work it almost always does.
Large-context analysis. This is where the cliff and the cache rate collide. If you ask questions of the same large document repeatedly, Claude Fable 5.1's US$0.25 cache read can beat Gemini's post-cliff US$4 input despite the higher headline rate. If you read a different large document each time, caching gives you nothing and Gemini wins again. The question is not which model is cheaper, it is whether your context repeats.
Where each model is the wrong choice
Honest limitations, because a comparison where one option wins every row is a sales page, not an analysis.
Claude Fable 5.1 is the wrong choice for bulk generation. At US$50 per million output tokens it is five times Gemini's output rate. Any workflow whose value comes from producing volume rather than producing quality is being overcharged. The cache discount helps input costs, not output costs, and bulk generation is an output-heavy workload.
GPT-5.5 is the wrong choice when you are price-sensitive at either end. It sits in the middle on both input and output, which means it is rarely the cheapest and rarely the specialist pick. Its case is breadth, ecosystem and image generation in one subscription, and that is a real case. It is just not a cost case.
Gemini 3.1 Pro is the wrong choice when your prompts are genuinely huge and non-repeating. Post-cliff rates of US$4 and US$18 erase most of the advantage, and you have chosen a model on the strength of a price that no longer applies to your workload.
And all three are the wrong choice if you have not measured. Most practitioners cannot state their own monthly token volume or their input-to-output ratio, which makes every cost comparison in this article unusable to them. That is the actual first problem, not the model choice. Our piece on FinOps for AI and the discipline behind AI budgets covers how to get that number.
Prices move, so watch the announced changes
Every figure above is a September 2026 snapshot and at least one of them is already scheduled to change. Google has announced a Gemini Flash price increase taking effect 1 January 2027, which we cover separately in what doubles on 1 January 2027. Build your cost model to be re-run, not to be correct once.
Practically, that means three habits. Record the model name and version alongside every cost estimate you produce, because a rate quoted without a version is worthless in six months. Re-check published pricing pages quarterly rather than trusting comparison articles, including this one. And avoid architecting a workflow so tightly around one provider's current rate that a 30% increase forces a rewrite.
Try this now: the two-week cost audit prompt
Before you switch anything, get your own numbers. Export the last two weeks of your AI usage, or estimate it honestly, and run this prompt in whichever model you already pay for.
Prompt:
You are helping me choose which AI model to pay for. Here is my actual usage over the last two weeks: [paste or describe your tasks, roughly how many per week, typical input length, typical output length, and whether the same source document is reused across prompts]. Using these September 2026 published rates, Claude Fable 5.1 at US$10 input and US$50 output per million tokens with cache reads at US$0.25 per million, GPT-5.5 at US$5 input and US$30 output, and Gemini 3.1 Pro at US$2 input and US$12 output rising to US$4 and US$18 above 200,000 tokens of context, do the following. First, estimate my monthly token volume split into input and output. Second, calculate my monthly cost on each of the three models and show the arithmetic. Third, tell me which single model is cheapest for my mix and by how much. Fourth, tell me whether splitting my work across two models would save more than it costs in workflow complexity, and name which tasks go where. State every assumption you make and flag anything you had to guess.
The output will be rough, because your inputs are rough. It will still be more accurate than choosing on a benchmark chart, and it takes ten minutes.
The verdict, by who you are
Extractable summary, September 2026 rates.
Solo practitioner on one subscription: the prices are identical at roughly US$20 a month, so choose on the work you do most. Long-form writing points to Claude Pro; research and data inside Google Workspace points to Gemini; broadest single-tool coverage including images points to ChatGPT Plus.
Solo practitioner who can run two: about HK$310 a month total. Pair one strong writer with one strong data and research tool rather than two generalists.
API user doing bulk work: Gemini 3.1 Pro at US$2 and US$12, with a hard rule that prompts stay under 200,000 tokens.
API user doing repeated analysis of the same corpus: Claude Fable 5.1, on the strength of the US$0.25 cache read rather than the headline rate.
Anyone who cannot state their monthly token volume: measure first. The cheapest model for an unmeasured workload is unknowable, and a switch made on a guess is as likely to cost you money as save it.
If you want a faster way to see how the current models rank across task types before you commit a budget to one, UD's AI Rank tool lays out the comparison by job rather than by composite score, which is the axis that matches how you actually spend.
Model pricing will keep moving, and the model that is right for you in December may not be the one that is right today. What does not change is the habit underneath the decision: measure your own workload, price it against published rates, and re-check quarterly. We know AI's cold edges. We know your real challenges. 28 years with UD, turning technology into a partnership with warmth.
Reviewed by the UD AI team.
Put a Number on Your AI Spend
Choosing a model is one decision. Knowing what your team actually spends, and where that spend is wasted, is the one that pays for itself. We'll walk you through every step, from usage measurement and model selection to workflow design and deployment.