Three Google image models die today. A new Gemini tier is cheap until 31 December. None of this is in the headlines, and all of it is in your next invoice.
The list price is now the least informative number a model vendor publishes. What determines your bill is the threshold you cross, the tokens you generate, the cache row nobody mentioned, and the date the introductory rate expires.
I have spent six weeks writing about capability, security, and governance. This week the news is almost entirely commercial fine print—and it is the most directly actionable cycle I have covered, because every item has a number attached, and one has a deadline that expires today.
Grok 4.6 and the 200,000-Token Toll Booth
SpaceXAI released Grok 4.6 on 12 August, thirty-five days after Grok 4.5. It scores 61 on the Artificial Analysis Intelligence Index — level with GPT-5.6 Sol, with Claude Opus 5 leading at 63 — and keeps the 500,000-token context window and the $2 per million input, $6 per million output headline rate.
On those numbers it looks like frontier parity at a substantial discount, and much of the coverage has said so.
Then read the pricing table.
xAI's documentation places Grok 4.6 in a separate long-context band at 200,000 prompt tokens. Once a prompt reaches that threshold, every token in the request bills at the doubled rate of $4 and $12—not the overage, the whole thing.
The arithmetic is worth doing slowly, because the shape is not intuitive. A request with 100,000 input and 10,000 output tokens costs roughly $0.26. A request with 250,000 input and 10,000 output costs roughly $1.12.
The second prompt is two and a half times larger. The bill is more than four times higher.
Crossing the line by a single thousand tokens is enough. A 201,000-token prompt pays the long-context rate on all 201,000.
A 500,000-token context window with a toll booth at 200,000 is a different product from one without. The window is what they advertise. The threshold is what you pay.
The Row That Was Not in the Launch Post
There is a second change that got no announcement at all. Cached input rose from $0.30 to $0.50 per million on the standard tier — a 67% increase — and was not mentioned in the launch materials.
For a long-running agent, cached input is where a large share of the spend actually lives. Stable system prompts, tool definitions and retrieved context get re-read on every step. A 67% rise on that row moves real money for exactly the workload this model is marketed at.
The Number That Reframes Everything
Now the figure that should be the headline and is nowhere near it.
Artificial Analysis measures the cost to complete its standard task set. Grok 4.6 comes in at $0.837 against Grok 4.5's $0.360 — a 2.32x increase, generation over generation, driven by roughly 47% more output tokens plus the cache price rise.
Same headline price. Same vendor. Successor model. More than double the cost to do the same work.
This is the clearest demonstration I have seen of something I keep arguing: per-token price tells you almost nothing. A model that reasons more verbosely at an identical rate is a more expensive model, and the only figure that captures it is cost per completed task against your own workload.
One reviewer's summary of the practical discipline is exactly right: price the model against your own prompt-size distribution, not the vendor's headline.
Two further caveats, stated fairly. Grok 4.6 has genuine strengths — step efficiency on long-horizon agent tasks is measurably better, and xAI's parity claim was measured at Grok's default effort against rivals' maximum tiers, which is a real result and not a leadership claim. And there is no EU region, which for some readers ends the conversation before pricing does.
Three Google Models Die Today
The second item has a deadline that expires while you are reading this.
Google is retiring three Imagen 4 model IDs on 17 August 2026 — the generate, ultra-generate and fast-generate variants — and directing developers to migrate to gemini-3.1-flash-image. The deprecations page flags today as the earliest possible shutdown date.
The migration is not a drop-in swap. The generate_images() method is gone entirely; image generation now runs through generate_content(). Teams still calling Imagen 4 endpoints need to re-test prompt adherence, aspect ratios, provenance-mark handling, latency and quotas before pipelines start failing.
That is a code change, a behavioural re-test and a quota re-check, on a date set by someone else, with no negotiation.
I published an argument five days ago that switching cost — measured in working days to replace a dependency — should be the primary metric shaping AI architecture, ahead of capability and price. I did not expect it to be tested this quickly, or this literally. Any team with Imagen 4 in a production pipeline is finding out its number today, empirically, whether or not it ever wrote one down.
Gemini 3.7 Flash, and the Cliff in December
Google also released Gemini 3.7 Flash, three weeks after 3.6 Flash. The capability jump is substantial on the published figures: FrontierCode 1.1 from 34.4% to 43.6%, DeepSWE v1.1 from 49% to 65.3%, and AutomationBench from 17% to 30.4%, with the 1M-token context window retained.
That AutomationBench movement — near enough a doubling — is the one worth watching if you build workflows rather than write code.
Introductory pricing runs at $0.75 per million input and $3.75 output through 31 December 2026, returning to $1.50/$7.50 in 2027.
So: cheap now, double in January, and the date is published in advance. That is the fourth dated pricing cliff I have had to put in front of you in six weeks, and it is worth pausing on how normal this has become.
The Pattern
Introductory pricing is now a standard commercial instrument in this market rather than an occasional promotion. Look at the pattern across a single quarter:
- Claude Sonnet 5 launched at an introductory rate with a September expiry — then the increase was cancelled and the rate made permanent.
- Claude Fable 5's grace period still ends on 30 September, moving to credits-only afterwards.
- GPT-5.6 Luna was cut 80% three weeks after launch, unprompted by any published schedule.
- Gemini 3.7 Flash is cheap until 31 December and doubles in 2027.
- Grok 4.6 kept its headline rate while raising the cache row and more than doubling cost per task.
Prices went up, down, sideways and away in the same quarter. Some changes were announced with dates, some were announced and then reversed, and at least two were not announced at all.
You cannot plan against this by choosing well. You can only plan against it by checking often and being able to move. Every dated cliff is a reminder that your pricing assumptions have an expiry date printed on them.
And It Is Already Obsolete
One more detail from the Grok launch that generalises past it.
Grok 4.6 arrived thirty-five days after Grok 4.5. Musk stated in late July that Grok 4.7, a larger model on a new architecture, would follow within weeks. So the model I have just spent a thousand words pricing may be superseded before your evaluation of it finishes.
The practical instruction that falls out of this is worth stating plainly: plan integrations around the model family, not the version string. If your code, prompts or contracts name a specific version, you have built a dependency with a shelf life measured in weeks rather than years.
That applies well beyond one vendor. Gemini went 3.6 Flash to 3.7 Flash in three weeks. The half-life of a current-flagship claim across this entire market is now shorter than most procurement cycles, which is not a complaint — it is a design constraint.
The Trust Dimension
One more item belongs here, because it bears on how much weight to give any vendor's framing.
The Future of Life Institute published its Summer 2026 AI Safety Index. No lab scored above C+. Anthropic led at C+, OpenAI, Google DeepMind at C, Meta at D+, and xAI, DeepSeek and Mistral at F. The report notes that the top four labs weakened their pause pledges, which it characterises as moving goalposts.
I want to handle this carefully, because two weeks ago I wrote approvingly about OpenAI pausing its Astra model over cyber capability it could not bound, and contrasted it with Anthropic having removed a pause commitment in February.
The FLI finding complicates that. If all four leading labs have weakened pause commitments over the period, then a single high-profile pause sits inside a broader pattern of loosening rather than standing against it. Both things can be true — the Astra decision looks genuinely creditable, and the aggregate direction of travel is the other way.
Note also that FLI is an advocacy organisation with a stated position on AI risk, and index methodologies involve judgment calls that reasonable people contest. Treat the grades as one informed assessment rather than an audit. The specific factual claim — that pause pledges were weakened — is the part that matters and the part that is checkable.
The operator takeaway is narrow: vendor safety framing is marketing until independently assessed, and the independent assessment this quarter is not flattering to anyone.
What To Do This Week
1. Check Whether You Call Imagen 4 — Today
If any pipeline touches those three model IDs, it may already be failing. Migration requires a method change rather than a string swap, plus re-testing prompt adherence, aspect ratios and quotas.
This is the only item here with a deadline that has already arrived.
2. Measure Your Prompt-Size Distribution
Not your average. The distribution — specifically, what proportion of your requests land above 200,000 tokens if you use or are considering Grok 4.6.
A workload sitting at 40,000 tokens sees one product. A workload packing 250,000 tokens per request is buying a different one at double the rate, and the only place that is written down is the pricing table. The same discipline applies to every provider with tiered thresholds.
3. Switch Your Metric From Price to Cost Per Task
The 2.32x figure exists because output verbosity and cache rates moved while the headline rate did not. No per-token comparison would have surfaced that.
Run your own version: take a representative workload, run it end to end on each candidate model, and record total spend and whether the output shipped without correction. That second measure is your Automation Ratio, and cost per acceptable output is the only figure that combines both.
4. Put Every Dated Cliff in One Calendar
Fable 5 on 30 September. Gemini 3.7 Flash on 31 December. Whatever your other providers have published. One calendar, owned by one person, reviewed monthly.
This is unglamorous and it is the difference between a planned migration and an emergency one. I laid out an earlier version of the pricing calendar in July and half of it has already changed, which is the argument for maintaining your own rather than relying on anyone's summary — including mine.
5. Practise Context Hygiene Regardless of Vendor
The practical response to threshold pricing is not avoiding long tasks. It is keeping tool results trimmed, avoiding whole-file pastes where a targeted read will do, and leaning on cached input for stable prompt sections.
A well-scoped tool call that returns exactly the resource an agent asked for is cheaper than a broad dump the model has to read past — and it is cheaper on every provider, threshold or not. This is task-model matching applied to context rather than to model choice.
The Broader Read
Six weeks ago I would have told you the interesting questions in AI were about capability. I no longer think that is where the decisions are.
Capability is abundant, converging and cheap. Four models now sit within a couple of points of each other on the main intelligence index, at prices that vary by more than they do. What separates outcomes is not which one you pick. It is whether you can read a pricing table, model your own distribution against it, notice a deprecation before it fires, and move when the answer changes.
That is the platform war as it actually plays out for a business: not a benchmark race, but an accumulation of thresholds, expiry dates, deprecation notices and unannounced cache-row changes, each individually small and collectively decisive.
It is also, I think, a large part of why 80% of executives report no measurable AI ROI. The returns are real and they are being eaten in the fine print by organisations that never assigned anyone to read it. A 2.32x cost increase that arrives with an unchanged headline price does not show up in any dashboard until the invoice does.
None of this requires a new vendor, a new model or a budget approval. It requires somebody whose job it is to know the numbers. That is a systems problem, it has always been the answer, and this week it comes with a deadline that expires today.
Frequently Asked Questions
Is Grok 4.6 cheap or not?
It depends entirely on your prompt sizes, which is the honest answer and the reason the question keeps getting answered wrongly. Below 200,000 tokens the $2/$6 rate is genuinely competitive against models it matches on the intelligence index. Above that threshold the whole request reprices at $4/$12. And on Artificial Analysis's standard task set it costs 2.32 times what Grok 4.5 did, because it generates more output tokens.
Does the doubled rate apply only to tokens above 200K?
No — and this is the detail most likely to surprise you on a bill. The higher rate applies to every token in the request. A 201,000-token prompt pays the long-context rate on all 201,000, not on the 1,000 above the line.
What do I do about the Imagen 4 retirement?
Check today whether anything you run calls those three model IDs. If so, migrate to the replacement model, and budget for re-testing rather than a straight swap — the image generation method itself changed, so prompt adherence, aspect ratios, provenance handling, latency and quotas all need verifying before you trust the pipeline.
Should I switch to Gemini 3.7 Flash while it is cheap?
Test it on your own workload before deciding anything, and go in knowing the rate doubles on 1 January 2027. Introductory pricing is a real saving and it is also a scheduled future cost increase. Both facts belong in the same decision, and the benchmark gains — particularly on automation tasks — are large enough to justify the test regardless.
How seriously should I take the FLI safety grades?
As one informed assessment rather than an audit. FLI is an advocacy organisation with a position, and index methodologies involve contestable judgment. The checkable factual claim — that the leading labs weakened pause commitments over the period — is the part worth carrying, and it is a useful corrective to any single vendor's safety messaging including the ones I have praised.
What is the single most useful action here?
Build one calendar of every published pricing change and deprecation date across your providers, and give one person responsibility for it. Every item in this article was published in advance. The businesses that get hurt are not the ones that were surprised; they are the ones where nobody was looking.
Related Reading
Stop Chasing the Biggest Model — task-model matching, and why cost per completed task beats per-token price every time.
The Automation Ratio Is the Only Metric That Predicts Survival — the quality half of the equation that cost per acceptable output depends on.
80% of Executives Report No Measurable AI ROI — the returns that get eaten in fine print nobody was assigned to read.
Fable 5's Return and the Sonnet 5 Pricing Calendar — an earlier version of the dated-cliff calendar, half of which has already changed.
The 'AI Business' Advice Is Wrong: Tool Delivery vs Process Ownership — why a business built on a tool collection feels every one of these changes and a process-owner absorbs them.
The AI Agent Platform War — the competitive dynamic that produces a new threshold, expiry or deprecation every fortnight.
The No-Code Automation Workflow Guide — building workflows with context hygiene and model choice as configuration rather than assumption.
About the Author
Hamza Baig is the founder of Hexona Systems, an AI automation agency serving clients across six continents, and the AI Automation Institute, a community of more than 40,000 entrepreneurs building with AI.
He has been featured in the GHL Top 50, Yahoo Finance and Brainz Magazine, and writes regularly on automation architecture, agent governance and the operational realities of AI deployment.
Read more analysis on the Hamza Automates blog, or get in touch to discuss an automation build.
Follow @hamza_automates on Instagram for daily automation breakdowns.
Note: pricing and deprecation details reflect published vendor documentation and contemporaneous coverage as of 17 August 2026 and change frequently — verify against provider pricing pages before making commercial decisions. Benchmark figures at launch originate with the vendors or with Artificial Analysis as indicated. This is not investment or legal advice.
About
Hamza Baig is the founder of Hexona Systems—an automation agency and softwareplatform that helps thousands of entrepreneurs and business owners implement AI-powered workflows at scale.








