Google released Gemini 3.8 Flash on 2 September 2026, three weeks after 3.7 Flash. Alongside it came Gemini 3.8 Flash Cyber, a separate model built for finding and patching security vulnerabilities.
Most coverage of a launch like this repeats the three benchmarks the press release highlights. We are going to do the opposite: put all fifteen published benchmarks in one table, next to the five models Google compared against, and then spend just as long on the rows where Gemini loses. A model you are about to put in production deserves both halves of the picture.

Gemini 3.8 Flash Benchmark Scores in Full
Here is the complete published table. Bold marks the best score in each row.
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| Input price, 0.75** | **5.00 | 4.00 | $2.00 | |||
| Output price, 3.75** | **25.00 | 20.00 | $12.00 | |||
| DeepSWE v1.1 | 73.7% | 65.3% | 74.0% | 53.8% | 72.7% | 69.6% |
| GDPVal-AA v2 (Elo) | 1545 | 1482 | 1824 | 1584 | 1710 | 1528 |
| Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.9% | 53.8% | 54.4% |
| Harvey's Legal Agent Benchmark | 10.0% | 8.8% | 6.7% | 5.0% | 2.5% | 0.8% |
| Terminal-bench 2.1 | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| Terminal-bench 4.0 | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
| GDP.PDF | 35.0% | 34.0% | 37.0% | 28.0% | 40.0% | 29.0% |
| CharXiv Reasoning | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
| LVBench | 87.8% agentic / 87.1% static | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
| HLE-Verified | 54.9% | 53.6% | 54.4% | 31.0% | 54.5% | 51.1% |
| OSWorld-2.0 | 59.0% | 50.6% | 75.4% | 42.6% | 62.6% | 50.2% |
| BioMysteryBench (human-solvable) | 88.8% | 87.1% | 90.1% | 87.5% | 79.5% | 83.8% |
| BioMysteryBench (human-difficult) | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
| LABBench2 | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |

Count the bold and the story is clear: Gemini 3.8 Flash takes nine of the thirteen capability rows, at a fifth of Claude Opus 5's input price. But four rows go the other way, and two of those four are not small losses.
The Three Benchmarks Where Gemini 3.8 Flash Loses
Start here, because this is what decides whether the model fits your workload.
Terminal-bench 4.0 — 19.1% against Claude Opus 5's 51.8%. This is the widest gap in the entire table, and it is not close. Terminal-bench 4.0 measures general agent capability over longer, messier tasks. Gemini improved a lot over 3.7 Flash's 11.2%, but it is still roughly a third of Opus 5. If you are building an agent that has to grind through open-ended terminal work without supervision, this row matters more than every row above it.
OSWorld-2.0 — 59.0% against 75.4%. Agentic computer use: driving a desktop, clicking through real applications. Gemini gained nine points over 3.7 Flash, and still sits well behind Opus 5.
GDPVal-AA v2 — 1545 Elo against 1824. Broad knowledge work, scored as an Elo rating rather than a percentage. Opus 5 is comfortably ahead, and GPT-5.6 Sol at 1710 is ahead too.
One more worth naming: GDP.PDF at 35.0%, where GPT-5.6 Sol takes it at 40.0%. Expert PDF comprehension is a common production task, so if your pipeline is mostly document extraction, do not assume the headline wins carry over.
Where Gemini 3.8 Flash Wins Convincingly
The wins cluster in a way that tells you what Google optimised for.
Professional domain work. Vals Finance Agent v2 at 61.4% and Harvey's Legal Agent Benchmark at 10.0% are both first place, and the legal number is first by a wide margin — Claude Opus 5 manages 6.7% and GPT-5.6 Terra manages 0.8%. Note that 10.0% is an all-pass rate on complex legal workflows, so a low absolute number is expected; what matters is that it is roughly fifty percent higher than the next model.

Structured software engineering. DeepSWE v1.1 at 73.7% sits a third of a point behind Opus 5's 74.0% — effectively a tie, at a fifth of the price. Terminal-bench 2.1 at 89.4% is an outright win.

Reasoning and science. HLE-Verified at 54.9% edges out GPT-5.6 Sol's 54.5% and Opus 5's 54.4%. LABBench2 at 86.2% and BioMysteryBench human-difficult at 56.5% are both clear firsts, and the second of those is a seven-point jump over Opus 5.

Charts and long video. CharXiv Reasoning at 86.2% with no tools, and LVBench at 87.8% agentic, both first. If your work involves reading figures or watching long footage, this is the strongest part of the model.

Gemini 3.8 Flash Pricing
3.75 per million output tokens, with no caching.
Read the footnote, though, because it changes the maths. That is an introductory price and it expires on 31 December 2026. From 1 January 2027 the rate becomes 7.50 per million output — exactly double.
Even after the doubling it stays cheaper than Claude Opus 5 at 25.00 and GPT-5.6 Sol at 20.00. But if you are building a cost model for next year, use the January numbers, not today's.
Gemini 3.8 Flash vs Gemini 3.7 Flash
Same price, three weeks apart, and the gains are real rather than cosmetic:
- DeepSWE v1.1: 65.3% → 73.7% (+8.4)
- Terminal-bench 4.0: 11.2% → 19.1% (+7.9)
- BioMysteryBench human-difficult: 43.5% → 56.5% (+13.0)
- OSWorld-2.0: 50.6% → 59.0% (+8.4)
- HLE-Verified: 53.6% → 54.9% (+1.3)
The biology jump of thirteen points is the largest. Since pricing is identical, there is no reason to stay on 3.7 Flash.
Prompt Injection Robustness
Google published a Gray Swan IPI result, which measures attack success rate within k attempts — lower is better.
At ASR@15, Gemini 3.8 Flash sits at 5.5%, and Gemini 3.8 Flash Cyber at 6.0%. For comparison, Gemini 3.7 Flash was at 9.2%, Claude Opus 5 leads at 4.8%, and the weakest model on the chart, DeepSeek V4 Pro, sits at 60.1%.
Cutting your own previous version's attack rate roughly in half is the single most useful number here for anyone shipping an agent that reads untrusted web pages or documents.

What Is Gemini 3.8 Flash Cyber?
A separate model for defensive security: finding vulnerabilities and writing patches. It runs with a more permissive set of mitigations than the general model, because a model that refuses to discuss exploit code is useless to a security team.
CyberGym Pass@1, vulnerability discovery in C/C++: 86.2%. That beats Gemini 3.5 Flash Cyber at 77.5%, GPT-5.5-Cyber at 85.6%, Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

Real-world vulnerability discovery across 20 programming languages: 71.0%, against Gemini 3.7 Flash at 58.9% and Gemini 3.5 Flash Cyber at 46.6%. That is a jump of over twenty-four points in two versions.

CWE-Bench Pass@1: 47.2%, just behind the leading frontier model at 47.8%. The interesting part is the cost axis: Google plots accuracy against average cost per rollout, and Gemini 3.8 Flash Cyber reaches roughly the same accuracy as Fable 5 at about a third of the cost per rollout.

The catch is access. Flash Cyber is not on the public API. It is released through the Fairwind Program, limited to trusted government authorities, critical infrastructure operators, and software maintainers. For most readers this is a model to know about, not one to plan around.
Where You Can Use Gemini 3.8 Flash Today
- Google Antigravity
- Gemini API — via Google AI Studio and Android Studio
- Gemini Enterprise
- Gemini app — Pro and Ultra subscribers
- AI Mode in Google Search
- Google Sheets
Gemini 3.8 Flash Cyber is available only through the Fairwind Program.

Should You Switch?
Three honest cases.
Switch, if you are on Gemini 3.7 Flash. Same price, better on every published row.
Switch, if your workload is document reasoning, financial or legal analysis, chart interpretation, long video, or structured coding, and you are currently paying Opus 5 prices. You get near-parity on DeepSWE and better numbers on the professional benchmarks for roughly a fifth of the input cost.
Do not switch, if your product depends on long-horizon autonomous agents or computer use. Terminal-bench 4.0 at 19.1% against 51.8%, and OSWorld-2.0 at 59.0% against 75.4%, are gaps that pricing does not close. Run your own evaluation on your own tasks before moving anything.
Frequently Asked Questions
What is the best Gemini 3.8 Flash benchmark score? Terminal-bench 2.1 at 89.4%, which is the highest absolute number in the table and also first place. On margin of victory, Harvey's Legal Agent Benchmark at 10.0% is the biggest win — roughly fifty percent above the next model.
Is Gemini 3.8 Flash better than Claude Opus 5? On nine of thirteen capability benchmarks, yes, at a fraction of the price. On Terminal-bench 4.0, OSWorld-2.0, GDPVal-AA v2 and BioMysteryBench human-solvable, no — and the first two are not close.
How much does Gemini 3.8 Flash cost? 3.75 per million output tokens until 31 December 2026, then 7.50 from 1 January 2027.
Is Gemini 3.8 Flash Cyber publicly available? No. It is distributed through the Fairwind Program to trusted government authorities, critical infrastructure operators and software maintainers.
What is HLE-Verified? A multidisciplinary expert reasoning benchmark spanning STEM and the humanities. Gemini 3.8 Flash scores 54.9%, the highest of the six models compared.
Recap
Gemini 3.8 Flash wins most of the published table and does it at $0.75 per million input tokens, which is a fifth of Claude Opus 5. The strongest results are in professional domains — finance, law, biology — plus chart reasoning and long video, and its prompt injection robustness roughly doubled over 3.7 Flash.
The losses are concentrated and worth respecting. Long-horizon agent work on Terminal-bench 4.0 and computer use on OSWorld-2.0 both sit far behind Opus 5, and GPT-5.6 Sol still leads on expert PDF comprehension. Remember too that today's price doubles on 1 January 2027.
Gemini 3.8 Flash Cyber is the more unusual release: 86.2% on CyberGym and 71.0% on real-world vulnerability discovery across twenty languages, but locked behind the Fairwind Program, so for most teams it sets expectations rather than plans.
Sources: Google's announcement and the published evaluation charts.