Google Gemini 3.8 Flash: What's New?
Gemini 3.8 Flash shipped 2 September — Google’s third Flash release in six weeks. Benchmark gains are modest, but the introductory price doubles on 1 January 2027, and a second model ships that you probably cannot use.
Advertisement

Google shipped Gemini 3.8 Flash on 2 September 2026 — its third Flash release in six weeks. The generation before it, 3.6 Flash, was still losing to Claude Opus 5 on most head-to-heads. It is faster to say what did not change: the price, the speed and the model class. Everything else moved.
Two things in this release matter more than the benchmark bumps. One is a scheduled price doubling that most coverage has skipped past. The other is a second model shipped alongside it, Gemini 3.8 Flash Cyber, which you almost certainly cannot use.
Everything below comes from Google's own announcement, model documentation and DeepMind benchmark pages, captured while writing.
What Actually Changed
Google frames 3.8 Flash as improving on 3.7 across software engineering, agentic tasks and multi-step reasoning in specialised domains. The behavioural description is the most useful part of the announcement:
The models "exhibit greater diligence — executing extra reasoning steps, and calling tools iteratively," especially on complex tasks.
That is a meaningful characterisation rather than marketing. A model that runs more reasoning steps and loops tools more persistently behaves differently in an agent harness: it finishes more tasks unattended, and it costs more tokens per task. If you are running 3.7 Flash on a fixed token budget, expect 3.8 to spend more of it per call — the same effect that makes agentic token spend hard to forecast on any model.
Benchmarks against 3.7 Flash
| Benchmark | 3.8 Flash | 3.7 Flash | Gain |
|---|---|---|---|
| HLE-Verified (expert reasoning) | 54.9% | 53.6% | +1.3pp |
| Vals Finance Agent v2 | 61.4% | 59.0% | +2.4pp |
| Harvey Legal Agent | 10.0% | 8.8% | +1.2pp |
Google also cites DeepSWE v1.1, where it says 3.8 Flash "outperforms most larger frontier models in autonomously solving complex engineering problems end to end" — though the published chart does not carry exact figures.
Read those gains honestly: they are real but incremental, in the 1–2.5 point range. This is a refinement release, not a generational jump. What makes it interesting is that the gains arrive at the same price and speed as 3.7 — Google is holding cost flat while pushing quality, which over three releases in six weeks compounds into something significant.
The Price Doubles on 1 January 2027
This is the detail to plan around. Gemini 3.8 Flash launched at:
| Period | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Now → 31 December 2026 | $0.75 | $3.75 |
| From 1 January 2027 | $1.50 | $7.50 |
That is a 100% increase, already scheduled and published. The introductory rate matches 3.7 Flash, so anyone budgeting from today's invoice is modelling roughly four months of pricing that expires.
If you are running Gemini Flash at volume, build your 2027 forecast on $1.50 / $7.50 and treat the next four months as a discount rather than a baseline. At the post-January rate Flash is still cheap — but it is not the same trade against Claude Opus 5 and GPT-5.6, which is a comparison worth re-running once the discount lapses.
Specs and Capabilities
| Property | Value |
|---|---|
| Model ID | gemini-3.8-flash |
| Input context | 1,048,576 tokens (~1M) |
| Max output | 65,536 tokens (64K) |
| Input modalities | Text, image, video, audio, PDF |
| Output | Text only |
| Thinking levels | low, medium, high |
Supported: caching, code execution, file search, function calling, structured outputs, search grounding, Google Maps grounding, URL context, batch and Flex inference, and computer use in preview.
Not supported: image generation, audio generation, and the Live API. Flash is a reasoning and tool-use model — for generative media you still reach for a different model.
One change that will break existing code
Thinking supports low, medium and high — but not minimal. If you have code passing a minimal thinking level to an earlier Flash model and you swap the model ID to 3.8, that call will not behave as before. It is a small change with a sharp edge, and it is the first thing to check when migrating.
Gemini 3.8 Flash Cyber
Shipped alongside 3.8 Flash is a cybersecurity-specialised variant, and it is the more unusual release of the two.
Google reports it surpasses 3.5 Flash Cyber and "significantly larger frontier models" on CyberGym, and that internal testing found real-world vulnerability discovery exceeding 70% across 20 programming languages.
Worth noting what Google reports honestly: on CWE-Bench patching, 3.8 Flash Cyber scores 47.2% against a leading frontier model's 47.8% — a narrow loss, published rather than omitted. That is a more credible presentation than a page of clean sweeps, and it tells you patching remains genuinely hard for every model.
You probably cannot access it. Flash Cyber is gated behind the new Fairwind Program, a restricted-access scheme for "trusted defenders" — government and critical infrastructure — by application. This is offence-defence asymmetry management: a model good at finding vulnerabilities is equally good at finding them for the wrong people. Treat announcements about it as industry news, not a product you can adopt this quarter.
Where You Can Use It
| Surface | Access |
|---|---|
| Gemini API | Generally available |
| Google AI Studio | Generally available |
| Android Studio, Stitch | Generally available |
| Gemini app | Google AI Pro and Ultra subscribers |
| AI Mode in Google Search | Pro and Ultra subscribers |
| Gemini in Google Sheets | Pro and Ultra subscribers |
| Gemini Enterprise | Available |
| 3.8 Flash Cyber | Fairwind Program only, by application |
Three Releases in Six Weeks
3.6 → 3.7 → 3.8 in six weeks is the story underneath the story, and it cuts both ways.
In your favour: quality improves without a price rise, and you inherit the gains by changing a model string. Against you: an eval suite validated against 3.7 is stale within weeks, prompts tuned to one model's diligence behave differently on the next — a prompt library built for 3.6 is not automatically valid on 3.8, and "latest" is a moving target in production.
The practical response is to pin explicit model IDs rather than aliases, keep a small regression suite you can rerun in an afternoon, and schedule model upgrades deliberately instead of drifting onto whatever is newest. A model that runs extra reasoning steps and loops tools more is exactly the kind of change that passes a smoke test and shifts your token bill.
Should You Upgrade from 3.7?
| If you… | Verdict |
|---|---|
| Run agentic workflows or tool loops | Yes — diligence gains land hardest here |
| Do software engineering tasks | Yes — the headline improvement area |
| Work in finance or legal analysis | Yes — the two benchmarks Google leads with |
| Run simple classification or extraction | Optional — 1–2pp gains will not show |
| Are on a tight per-task token budget | Test first — extra reasoning steps cost tokens |
| Pass a minimal thinking level | Fix that call before switching |
| Are budgeting 2027 spend | Model $1.50 / $7.50, and re-check it against the alternatives |
The Verdict
Gemini 3.8 Flash is a refinement release that is worth taking, with one asterisk. The benchmark gains are modest — 1 to 2.5 points — but they arrive at unchanged price and speed, and the behavioural shift toward more diligent tool use is the kind of change that matters more in an agent loop than a leaderboard suggests.
The asterisk is the pricing. Introductory rates through 31 December 2026, doubling on 1 January 2027, is a substantial change to schedule into any forecast that outlives this year. It does not make Flash expensive — it makes it a different value proposition against the alternatives, which is exactly the comparison worth running before you commit.
Keep Reading
Claude Opus 5 vs Gemini 3.6 Flash covers how the previous generation stacked up, and Claude pricing explained breaks down the subscription side of the frontier tier. Or browse all guides and prompts on PromptsRush.
Frequently Asked Questions
10 questions answered

