Gemini Flash Is Half Price Until 31 December. Then It Is Not
Three of the current Flash models carry a promotional rate that expires at the end of the year, and the pricing page says so in the same row as the price. If you are costing a project on today's number, you are costing it on a number with a deadline.

Quick answer
Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash all bill at $0.75 per million input tokens and $3.75 per million output through 31 December 2026, and $1.50 / $7.50 after that — a straight doubling on a fixed date. The odd part is that Gemini 3.5 Flash, the older model, already costs $1.50 / $9.00, so the newest Flash is currently the cheapest and will still be no worse than the one it replaced. If you want a rate with no expiry attached, 3.5 Flash-Lite is $0.30 / $2.50. Budget on the post-January number and treat the discount as a windfall, not a baseline.
Model pricing pages are written to be read once, at the moment you are choosing, and never again. That is how a promotional rate becomes a budget: you check a number in September, build against it, and find out in January that the number had a date attached and you did not read it.
Gemini's Flash tier has exactly that shape right now. Three models share one promotional rate, and the rate expires at the end of the year.
To be clear about what this is: a pricing comparison, not a usage report. I have not run a production workload on Gemini Flash, so there is no bill of mine in here. Everything below comes from Google's pricing page as read on 24 September 2026, plus arithmetic you can check.
What the change does to a monthly bill
Take an example month of 50 million input tokens and 10 million output tokens on Gemini 3.8 Flash. That is an illustration, not a measurement. At today's rate it costs $37.50 for input and $37.50 for output: $75 in total. From 1 January 2027 the same month costs $75 for input and $75 for output: $150. Same work, same model, twice the bill.
Change the mix and the story changes. Swap the ratio to 10 million in and 50 million out, which is closer to an agent that writes more than it reads, and the month goes from $195 to $390. The doubling is the same percentage either way, but the output price is five times the input price, so output-heavy work is where the dollar amount grows fastest.
The rates, and which ones have a deadline
| Model | Input / 1M | Output / 1M | Expires |
|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | 31 Dec 2026, then $1.50 / $7.50 |
| Gemini 3.7 Flash | $0.75 | $3.75 | 31 Dec 2026, then $1.50 / $7.50 |
| Gemini 3.6 Flash | $0.75 | $3.75 | 31 Dec 2026, then $1.50 / $7.50 |
| Gemini 3.5 Flash | $1.50 | $9.00 | No expiry listed |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | No expiry listed |
Read the last column before the first. Three models are on a countdown and two are not, and the two that are not include the cheapest option on the page.
Why the newest model is the cheap one
The usual assumption is that a newer model costs more, so you stay on the old one to save money. Here that is backwards. Gemini 3.8 Flash is less than half the output price of Gemini 3.5 Flash today, and after January it is still a fifth cheaper. Whatever reason you might have for staying on 3.5 Flash, cost is not it.
What this tells you about the pricing is that the promotional rate is not a discount on 3.8 Flash so much as a repricing of the whole Flash tier, with the old model left at its old number. The older 2.5 Flash sits lower again, at $0.30 / $2.50 for text, so a newer model at a higher price than its predecessor is not a rule here in either direction.
Where a doubling actually hurts
Input and output do not rise by the same amount in practice, because most workloads are lopsided. A summarising job reads a lot and writes a little, so it lives on the input price. An agent that plans, calls tools and explains itself writes far more than it reads, and it lives on the output price — which is the one going from $3.75 to $7.50.
That is the number to check against your own billing page. If your output volume is small, January is a non-event. If you are running something that talks to itself a lot, it is the whole cost of the project changing on a date you did not choose. The same asymmetry shows up when running models on your own hardware starts to look reasonable: the break-even is set by output tokens, not by how clever the model is.
What to do before January
Take your last full month of usage and multiply it by the post-expiry rate. That is the only calculation that matters, and it takes five minutes. If the answer is small, do nothing and enjoy the discount until it ends. If the answer is large, you have three months to test 3.5 Flash-Lite at $0.30 / $2.50 on the same work and find out whether the cheap model was good enough all along — which, for the routine half of most pipelines, it usually is.
What you should not do is budget the project at $0.75 and put the date out of your mind. The page tells you what happens next. Prices above are from Google's Gemini API pricing page as read on 24 September 2026.
Pros and cons
Pros
- The three newest Flash models are all at the same promotional rate, so moving between them costs nothing
- The expiry date is published rather than buried in a footnote you find later
- Gemini 3.8 Flash is cheaper today than the 3.5 Flash it supersedes, and still cheaper after the rise
- The free tier on AI Studio has no token charge at all, which makes testing genuinely free
Cons
- A doubling on a fixed date is a budget cliff, and nothing in the API warns you when you cross it
- Output tokens rise from $3.75 to $7.50, and output is where agent workloads spend most of their money
- Four Flash variants at three different prices is more choice than most projects need
- Promotional pricing sets an expectation that the post-January rate then reads as an increase
Alternatives worth considering
$0.30 in / $2.50 out per million, with no promotional expiry attached. The cheapest Gemini that is not on a countdown.
$0.30 in / $2.50 out for text, image and video; $1.00 for audio input. Older, and priced like it.
$1.25 in / $10.00 out up to 200k tokens, rising to $2.50 / $15.00 above that. The tiering by context length catches people out.
No token charge. The honest way to find out whether a cheaper model is good enough before you pay for the expensive one.
Frequently asked questions
What exactly happens on 1 January 2027?
When I read the pricing page on 24 September 2026, Gemini 3.8, 3.7 and 3.6 Flash were listed at $0.75 per million input and $3.75 per million output "through Dec 31, 2026", with $1.50 and $7.50 shown as the rate afterwards. Google publishes both numbers in the same row, so there is no ambiguity about the direction.
Is the newest Flash model the cheapest one?
Right now, yes, and that is the counter-intuitive bit. Gemini 3.8 Flash at $0.75 / $3.75 undercuts Gemini 3.5 Flash at $1.50 / $9.00. Even after the promotional rate ends, 3.8 Flash at $1.50 / $7.50 is still cheaper on output than the older model. There is no version of this where staying on 3.5 Flash saves you money.
Should I switch away from Gemini before January?
Only if your own numbers say so. A doubling sounds alarming and may be a rounding error on your bill, or it may be the largest line on it — that depends entirely on your output token volume, which is why the worked example above is an example and not a recommendation. Work out what your last full month would have cost at the new rate before you move anything.
Sources
Everything factual in this article traces back to one of these. Vendors change pricing and limits without changing the URL, so each entry records the date I last read it.
- Gemini API pricing
Googlechecked September 24, 2026
Written by
Parth Patel
Founder and editor
Parth runs ToolNest and is responsible for everything published here. He researches each article from vendor documentation, changelogs, pricing pages and published reporting, drafts with AI assistance, then checks and edits every claim before it goes live. Where he has not used a tool himself, the article says so.