Google is currently selling Gemini 3.7 Flash at US$0.75 per million input tokens and US$3.75 per million output tokens. On 1 January 2027, both figures double. If your AI agents run on this model, your inference line doubles on a date already in the calendar.
That is the whole decision. Everything below is how to price it, when to commit, and where the model genuinely is not the right answer.
How much does Gemini 3.7 Flash cost right now?
Through 31 December 2026, Gemini 3.7 Flash costs US$0.75 per million input tokens and US$3.75 per million output tokens, with context caching at US$0.075 per million tokens. Google launched the model on 13 August 2026 with an explicit 50% introductory discount against its standard rate.
The discount is not a promotion in the marketing sense. Google published the end date at launch, so this is a scheduled price change rather than an offer that might quietly extend. Treat it as a known future cost, not a risk.
For context on why the price matters more than it used to: an autonomous agent turns one user request into a long sequence of model calls, reasoning tokens and tool interactions. Per-token cost compounds in a way it never did for a chat interface where a human types once and reads once.
What exactly changes on 1 January 2027?
On 1 January 2027, input rises from US$0.75 to US$1.50 per million tokens, output rises from US$3.75 to US$7.50 per million tokens, and context caching rises from US$0.075 to US$0.15 per million tokens. Nothing about the model changes. Only the invoice does.
The post-discount rate is not a penalty rate. It matches Gemini 3.6 Flash's standard API pricing of US$1.50 and US$7.50, which is the tier Google intended for this model class all along. What you have until December is the discount, not the baseline.
The planning implication is specific. Any 2027 budget built on today's invoices will be short by roughly half on the inference line, and the shortfall lands in the same month your finance team closes the year. Budget from the January number, not the December one.
Gemini 3.7 Flash price list at a glance
Every figure below is Google's published API list price. Convert at your own treasury rate; the illustrative Hong Kong dollar figures in this article use HK$7.80 to US$1.
--- Input tokens, until 31 Dec 2026: US$0.75 per million.
--- Input tokens, from 1 Jan 2027: US$1.50 per million.
--- Output tokens, until 31 Dec 2026: US$3.75 per million.
--- Output tokens, from 1 Jan 2027: US$7.50 per million.
--- Context caching, until 31 Dec 2026: US$0.075 per million tokens.
--- Context caching, from 1 Jan 2027: US$0.15 per million tokens.
--- Effective change on 1 Jan 2027: 100% increase on input, output and caching.
--- Release date: 13 August 2026, three weeks after Gemini 3.6 Flash.
--- Enterprise access: Gemini API via Google AI Studio, Gemini Enterprise Agent Platform, and the Gemini Enterprise app.
How does that compare with Claude Sonnet 5 and GPT-5.6 Terra?
In Google's own published comparison table, Claude Sonnet 5 lists at US$2 per million input tokens and US$10 output, and GPT-5.6 Terra at US$2 input and US$12 output. Against those rates, Gemini 3.7 Flash today is roughly a third of the price. From January 2027 it is roughly three quarters.
That gap narrowing matters more than the headline. Today the price difference is large enough to justify accepting a weaker model on some tasks. In 2027, at US$1.50 and US$7.50 against US$2 and US$10, the argument becomes far closer, and capability differences start to decide it rather than price.
Two cautions before anyone builds a spreadsheet on those three numbers. Competitor rates are quoted from Google's own benchmark table, so verify them against each vendor's current published pricing before committing. And list price is not spend. A cheaper model that needs three attempts where a costlier one needs one is not cheaper.
The metric that decides this is cost per successfully completed task, not price per million tokens. That is a measurement discipline, and it belongs in the same practice as the rest of your AI FinOps work.
What does this actually cost a Hong Kong enterprise?
Take a customer-service agent at a 300-person Hong Kong professional services firm handling 40,000 conversations a month, averaging 12,000 input tokens and 2,000 output tokens per conversation once retrieval, tool calls and multi-turn context are counted. At today's rates that is US$660 a month.
The arithmetic, so you can substitute your own volumes:
--- Input: 40,000 conversations x 12,000 tokens = 480 million tokens. At US$0.75, that is US$360.
--- Output: 40,000 conversations x 2,000 tokens = 80 million tokens. At US$3.75, that is US$300.
--- Monthly total in 2026: US$660, roughly HK$5,150.
--- Monthly total from January 2027: US$1,320, roughly HK$10,300.
--- Annual difference: US$7,920, roughly HK$61,800 on one workload.
These are illustrative figures built from Google's list prices and stated assumptions, not a quote. Your token profile will differ, and long retrieved contexts move the input number fastest.
Note the shape of the result. HK$62,000 a year on a single agent is not a board-level number. Multiply it across four or five production agents, add the second and third year, and it becomes one. This is why the price change is a planning problem rather than a crisis: it is entirely predictable, and entirely ignorable until it is not.
The comparison worth running internally is against your per-seat AI licensing. Token pricing scales with usage; seat pricing scales with headcount. Most Hong Kong enterprises now carry both, and the two behave very differently as adoption grows. Per-seat Copilot economics is the other half of the same budget conversation.
Where does Gemini 3.7 Flash actually lose?
Google's own benchmark table does not show Gemini 3.7 Flash beating higher-priced competitors everywhere. GPT-5.6 Terra leads on Terminal-bench 2.1 at 87.4% against 85.8%, on DeepSWE v1.1 at 69.6% against 65.3%, and on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads Agent's Last Exam at 33.3% against 26.3%.
Where it does win, it wins clearly. FrontierCode 1.1 Main at 43.6%, against 42.7% for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra. Code Arena Elo of 1588. AutomationBench, Google's enterprise workflow measure, at 30.4% against 23.6% for Terra and 10.7% for Sonnet 5. Complex PDF comprehension at 34.0% against 28.0% and 24.7%.
Read that split honestly. Gemini 3.7 Flash is strong on document-heavy enterprise workflow automation and production code quality. It trails on long-horizon terminal and desktop operating tasks. If your agent lives inside documents and business systems, the price case is excellent. If it lives inside a terminal or drives a desktop, the cheaper token may cost you more in retries.
Two further limitations belong on the record. Artificial Analysis places Claude Opus 5 at 63 on its Intelligence Index against Google's reported 56 for Gemini 3.7 Flash, so this is not a frontier model. And Google has not shipped Gemini 3.5 Pro; its latest released general-purpose Pro model remains Gemini 3.1 Pro from February 2026. If your roadmap assumes a Google flagship arriving on schedule, that assumption currently has no date attached to it.
Should you migrate workloads now, or wait?
Migrate now if the workload is document-heavy, high-volume and already stable, because you capture four months of half-price inference and a measured baseline before rates change. Wait if the workload is terminal-based, agentic-desktop, or still changing weekly, because migration effort will outweigh the discount you can still collect.
A practical sequence for the remaining months of 2026:
--- Instrument first. If you cannot report tokens per completed task by workload, you cannot evaluate any of this.
--- Run one production workload on 3.7 Flash for two weeks and measure retries, not benchmarks.
--- Rebuild your 2027 inference budget at US$1.50 and US$7.50, then check whether the workload still clears its return threshold.
--- Use context caching deliberately. At US$0.075 today it is the cheapest lever available, and it doubles in January alongside everything else.
--- Keep one competing model qualified. Single-vendor inference is a pricing position, not an architecture.
The honest caveat on our side: if you have in-house engineering capacity, buying tokens directly and managing the workload yourself will always be cheaper than any managed arrangement. UD's AI Staff Solution and Cloud AI Staff exist for organisations that do not have that capacity, or do not want to spend it on model selection and cost engineering. If you do, keep it in-house and use the framework above.
What is the right next step?
Pull your last three months of inference invoices, split them by workload, and re-run the total at double the rate. If any workload stops making sense at the 2027 price, you have four months to redesign it, cache it, move it, or retire it. That is a comfortable amount of time and an uncomfortable amount of work.
What you should not do is discover this in January. A price change published five months in advance is the easiest kind of budget problem to solve and the most embarrassing kind to be surprised by.
Whether you handle it internally or bring in help, the work is the same: measure per task, model the 2027 number, and decide workload by workload. We understand AI. We understand you. With UD by your side, AI never feels cold.
Reviewed by the UD enterprise AI team.
Model your 2027 AI cost with us
Now that you have the numbers, the next step is applying them to your own workloads. We'll walk you through every step, from token instrumentation and per-task cost modelling to workload selection and deployment, backed by 28 years of Hong Kong enterprise experience.