the AI bench
VERIFIED SEPTEMBER 2026
All fast takes

FAST TAKE · 2026-09-02 · GEMINI 3.8 FLASH (+ 3.8 FLASH CYBER)

Gemini 3.8 Flash: same price as 3.7, three points smarter, and it thinks longer by design

Google released Gemini 3.8 Flash on September 2, 2026 — GA in the Gemini API from day one as `gemini-3.8-flash`, and its third Flash model in six weeks after 3.6 and 3.7. The price did not move: $0.75 in / $3.75 out per 1M tokens, the same introductory rate through December 31, 2026, with $1.50 / $7.50 published for January 1, 2027. What moved is the model: Google calls it "our most intelligent workhorse model", and an independent measurement agrees, at a cost — it uses about 30% more output tokens per task than 3.7 Flash.

Verdict: Google's third Flash in six weeks: measurably smarter than 3.7, same $0.75/$3.75, but it spends ~30% more tokens per task


The take

Google's claims first, labelled as such. From the launch post: "delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains"; "54.9% on HLE-Verified"; strongest on DeepSWE v1.1 (long-horizon software engineering) and on the Vals Finance Agent V2 and Harvey legal-agent benchmarks. Same 1M-token context and 64K max output as 3.7 Flash; `thinking_level` low / medium / high with medium the default, and the `minimal` level is gone. Google's own migration checklist also drops `temperature`, `top_p` and `top_k` for this model. The Antigravity managed agent now defaults to 3.8 Flash.

The independent read. Artificial Analysis (September 2): "Gemini 3.8 Flash (high) scores 59 on the Artificial Analysis Intelligence Index, up 3 points from Gemini 3.7 Flash (high, 56)" — level with GPT-5.6 Sol at xhigh effort and Grok 4.6 at medium, on their index. The catch is in the same report: cost per task is "$0.58 … up ~40% from Gemini 3.7 Flash ($0.40) despite unchanged per-token pricing, driven by a 30% increase in output tokens per task and more turns on agentic evaluations", and time per task rises from 2.2 to 2.5 minutes. Google says the same thing in its own words — "3.8 Flash works harder" — and keeps 3.7 Flash "fully supported for efficiency-first workloads".

What that means for our cost calculator: the Gemini mid-tier row is renamed 3.7 → 3.8 Flash and its per-million rate is unchanged at $2.25 blended. But a per-token price is not a per-task price. If your workload is agentic and you leave the default thinking level on, expect the monthly bill to run higher than the 3.7 figure for the same jobs; drop `thinking_level` to `low` or stay on 3.7 Flash if tokens are the constraint. We still do not pre-schedule the January 1 rise — a dated future price is exactly the shape that produced the cancelled Sonnet 5 bump — the guard test asserts $2.25 and re-quoting the vendor page in January is the only trigger.

Two side notes. Gemini 3.8 Flash Cyber ships alongside it, restricted to "trusted defenders" through Google's new Fairwind Program (governments, critical-infrastructure operators, software maintainers), with no public price — nothing for this site to price. And there is still no Gemini 3.x Pro above 3.1 Pro Preview at $2 / $12: every Google model page we read on September 7 lists 3.5, 3.6, 3.7 and 3.8 as Flash-tier only, so the seventh sweep in a row closes that question the same way. Closed weights, API and enterprise only; no local pick is touched.

Where this fits

Models: Qwen3.8-27B · DeepSeek V4-Flash · GLM-5.1

Hardware: NVIDIA RTX 5090 · Mac Studio M4 Max 64 GB

Sources

Next step

Try this in the planner