the AI bench
VERIFIED AUGUST 2026
All fast takes

FAST TAKE · 2026-08-13 · GOOGLE GEMINI 3.7 FLASH

Gemini 3.7 Flash lands at half price, and we were quoting you double

Google shipped Gemini 3.7 Flash on August 13 and put both it and the 3.6 Flash it replaces on introductory pricing at $0.75 in / $3.75 out per 1M tokens. We had been publishing $1.50 / $7.50 as the current rate. That was right when we recorded it on August 3 and wrong by the time you read it — and the error ran in the direction that makes local hardware look better than it is.

Verdict: A new Google workhorse at half the price — and a correction: we published the old, higher price for two weeks


The take

The new model first. `gemini-3.7-flash` is described by Google as "our most capable Flash model for agentic workflows and multimodal reasoning," and it takes over the workhorse tier from Gemini 3.6 Flash, which shipped only three weeks earlier on July 21. Same 1M context window. We treat this as a rename in the cost calculator rather than a third Google row, the same way we handled Opus 4.6 → 4.7 → 4.8 and Grok 4.5 → 4.6: the dropdown shows the current flagship, not a museum.

Now the correction, because it matters more than the model. Google's own pricing page lists, for both 3.7 and 3.6 Flash: input *"$0.75 through December 31, 2026. $1.50 starting January 1, 2027"* and output *"$3.75 through December 31, 2026. $7.50 starting January 1, 2027."* Google Cloud's Agent Platform page says the same thing in one sentence: *"Gemini 3.7 Flash and Gemini 3.6 Flash are offered with introductory pricing of $0.75 / $3.75 per 1M tokens input / output through December 31, 2026. Starting January 1, 2027, standard pricing of $1.5 / $7.5 per 1M tokens input / output will apply."* Our $1.50 / $7.50 was Gemini 3.6 Flash's launch price on July 21, correct on August 3 when we recorded it, and superseded when Google halved it.

The direction of the error is the part worth owning. Overstating a cloud price by 2× makes local hardware pay back faster than it really does, in a tool people use to decide whether to spend a few thousand dollars on a GPU. The honest correction is unflattering to the local case: Google's mid-tier is half what we said, so payback against it takes longer, not shorter. This is the second time in a month a cloud correction has run this way — the cancelled Sonnet 5 increase was the first.

One trap avoided and one carry-forward closed. The pricing page shows several tiers per model, and its Priority column for 3.7 Flash reads $1.35 / $6.75 while Batch reads $0.375 / $1.875; a first pass at the page nearly landed the Priority figures as standard, and the same confusion makes Gemini 3.1 Pro Preview appear to be $3.60 / $21.60. It is not — Pro is **unchanged** at $2 / $12 standard, and reading the wrong column would have introduced a fresh error while fixing an old one. Separately: for the sixth sweep running there is still no Gemini 3.5, 3.6 or 3.7 **Pro**. Enumerating every model heading on the pricing page returns only Gemini 3.1 Pro Preview, Gemini 2.5 Pro and Gemini 3 Pro Image. That generation remains Flash-tier only.

Finally, the January 1 date is a scheduled future price rise — exactly the shape that produced our worst recent error, when a stale "bump this on September 1" note in our own files instructed a later session to break a correct price after Anthropic cancelled the increase. So there is no reminder anywhere to bump this in January. Instead the guard test asserts $0.75 / $3.75 **and explicitly asserts it is not** the 2027 figure. If Google's page shows the higher rate in January, the test and the value change together, after checking the page — not before.

Where this fits

Models: Qwen3.8-27B · Qwen 3.6-35B-A3B

Hardware: NVIDIA RTX 5090 · NVIDIA RTX 4090 · Mac Studio M4 Max 64 GB

Sources

Next step

Try this in the planner