Gemini 3.8 Flash Benchmarks: Every Score Compared

Every Gemini 3.8 Flash benchmark score in one table, compared with Gemini 3.7 Flash, Claude Opus 5, Claude Sonnet 5 and GPT-5.6, plus pricing and where it loses.

Sep 3, 202613 min readFollow

Topics You Will Master

Every score from Google's launch comparison table, in one place
Where Gemini 3.8 Flash beats Claude Opus 5, and the three places it does not
What Gemini 3.8 Flash costs now, and what it will cost in January
What Gemini 3.8 Flash Cyber is and who can actually use it

Google released Gemini 3.8 Flash on 2 September 2026, three weeks after 3.7 Flash. Alongside it came Gemini 3.8 Flash Cyber, a separate model built for finding and patching security vulnerabilities.

Most coverage of a launch like this repeats the three benchmarks the press release highlights. We are going to do the opposite: put all fifteen published benchmarks in one table, next to the five models Google compared against, and then spend just as long on the rows where Gemini loses. A model you are about to put in production deserves both halves of the picture.

Bestseller

Master LangGraph and LangChain

Agentic RAG and Chatbot, AI Agent with LangChain v1, Qwen3, Gemma3, DeepSeek-R1, LLAMA 3.2, FAISS Vector Database

Enroll on Udemy 30 day refund, lifetime access

Gemini 3.8 Flash and Gemini 3.8 Flash Cyber


The Short Answer

Gemini 3.8 Flash leads 8 of the 14 capability benchmarks in Google's launch comparison, at 15% of Claude Opus 5's input price. It is strongest on professional-domain work — finance, law, biology — plus chart reading and long video.

It is weakest exactly where autonomous agents live. On Terminal-bench 4.0 it scores 19.1% against Opus 5's 51.8%, and on computer use (OSWorld-2.0) 59.0% against 75.4%. Price does not close those two gaps.


Advertisement

Gemini 3.8 Flash Benchmark Scores in Full

Here is Google's full launch comparison table. Bold marks the best score in each row. Google has published other evaluations separately, so this is that comparison in full, not every number that exists for the model.

Benchmark Gemini 3.8 Flash Gemini 3.7 Flash Claude Opus 5 Claude Sonnet 5 GPT-5.6 Sol GPT-5.6 Terra
Input price, $/1M tokens $0.75 $0.75 $5.00 $2.00 $4.00 $2.00
Output price, $/1M tokens $3.75 $3.75 $25.00 $10.00 $20.00 $12.00
DeepSWE v1.1 73.7% 65.3% 74.0% 53.8% 72.7% 69.6%
GDPVal-AA v2 (Elo) 1545 1482 1824 1584 1710 1528
Vals Finance Agent v2 61.4% 59.0% 58.6% 53.9% 53.8% 54.4%
Harvey's Legal Agent Benchmark 10.0% 8.8% 6.7% 5.0% 2.5% 0.8%
Terminal-bench 2.1 89.4% 85.8% 89.1% 80.4% 88.8% 87.4%
Terminal-bench 4.0 19.1% 11.2% 51.8% 12.4% 37.3% 23.6%
GDP.PDF 35.0% 34.0% 37.0% 28.0% 40.0% 29.0%
CharXiv Reasoning 86.2% 84.5% 83.7% 70.1% 85.8% 85.9%
LVBench 87.8% agentic / 87.1% static 85.4% 75.4% 68.5% 82.1% 78.9%
HLE-Verified 54.9% 53.6% 54.4% 31.0% 54.5% 51.1%
OSWorld-2.0 59.0% 50.6% 75.4% 42.6% 62.6% 50.2%
BioMysteryBench (human-solvable) 88.8% 87.1% 90.1% 87.5% 79.5% 83.8%
BioMysteryBench (human-difficult) 56.5% 43.5% 49.4% 34.1% 44.7% 49.4%
LABBench2 86.2% 82.1% 84.2% 80.1% 82.1% 81.2%

Count the bold and the story is clear: Gemini 3.8 Flash leads on eight of the fourteen capability rows, at 15% of Claude Opus 5's input price. Six rows go the other way, and two of those are not small losses.

Note

These rows come from different evaluation suites with different datasets, harnesses, metrics and agent configurations. They are not a single ranking, and a percentage on one benchmark is not comparable with a percentage on another. Treat small gaps with particular caution.


Advertisement

Where Gemini 3.8 Flash Loses

Start here, because this is what decides whether the model fits your workload. Six rows go against Gemini, though three of them are close enough to be noise: DeepSWE v1.1 (73.7% against 74.0%), BioMysteryBench human-solvable (88.8% against 90.1%) and GDP.PDF, covered below. The three that matter are these.

Terminal-bench 4.0 — 19.1% against 51.8%

This is the widest gap in the entire table, and it is not close. Terminal-bench 4.0 measures general agent capability over longer, messier tasks. Gemini improved a lot over 3.7 Flash's 11.2%, but it is still roughly a third of Opus 5. If you are building an agent that has to grind through open-ended terminal work without supervision, this row matters more than every row above it.

OSWorld-2.0 — 59.0% against 75.4%

Agentic computer use: driving a desktop, clicking through real applications. Gemini gained nine points over 3.7 Flash, and still sits well behind Opus 5.

GDPVal-AA v2 — 1545 Elo against 1824

Broad knowledge work, scored as an Elo rating rather than a percentage. Opus 5 is comfortably ahead, and GPT-5.6 Sol at 1710 is ahead too.

GDP.PDF — 35.0% against 40.0%

Expert PDF comprehension is a common production task, so if your pipeline is mostly document extraction, do not assume the headline wins carry over.


Advertisement

Where Gemini 3.8 Flash Wins

The wins cluster in a way that tells you what Google optimised for.

Vals Finance Agent v2 at 61.4% and Harvey's Legal Agent Benchmark at 10.0% are both first place, and the legal number leads by a wide relative margin — Claude Opus 5 manages 6.7%, GPT-5.6 Terra 0.8%.

Vals Finance Agent v2 scores: Gemini 3.8 Flash 61.4%, ahead of Gemini 3.7 Flash and frontier models

Vals Finance Agent v2. Gemini 3.8 Flash takes 61.4%, 2.4 points above 3.7 Flash and 2.8 above Opus 5.

Harvey's Legal Agent Benchmark: Gemini 3.8 Flash 10.0% all-pass rate versus 6.7% for Claude Opus 5

Harvey's Legal Agent Benchmark. 10.0% is an all-pass rate on complex legal workflows, so a low absolute number is expected — the point is that it is about 49% above the next model.

Reasoning and science

HLE-Verified at 54.9% is the highest reported score here, but only narrowly: GPT-5.6 Sol is at 54.5% and Opus 5 at 54.4%, a spread too small to call anyone better at expert reasoning. The clearer wins are LABBench2 at 86.2% and BioMysteryBench human-difficult at 56.5%, the latter seven points clear of Opus 5.

HLE-Verified scores: Gemini 3.8 Flash 54.9%, narrowly ahead of GPT-5.6 Sol and Claude Opus 5

HLE-Verified. The top three sit inside half a percentage point — read this as parity, not a lead.

Software engineering

DeepSWE v1.1 at 73.7% sits 0.3 points behind Opus 5's 74.0%, at 15% of the price. Terminal-bench 2.1 at 89.4% edges Opus 5's 89.1% — again by 0.3 points, so parity rather than a lead.

DeepSWE v1.1 long-horizon software engineering: Gemini 3.8 Flash 73.7% versus 65.3% for Gemini 3.7 Flash

DeepSWE v1.1. The meaningful comparison here is against 3.7 Flash — an 8.4-point gain — rather than the statistical tie with Opus 5.

Charts and long video

CharXiv Reasoning at 86.2% with no tools, and LVBench at 87.8% agentic / 87.1% static, are both first place by clear margins — LVBench by more than five points over GPT-5.6 Sol. Google published no standalone chart for either, so the table above is the only source for these two.

Gemini 3.8 Flash Pricing

$0.75 per million input tokens and $3.75 per million output tokens. Those are the standard uncached rates, which is what Google's table reports; cached input is priced separately.

Read the footnote, though, because it changes the maths. That is an introductory price and it expires on 31 December 2026. From 1 January 2027 the rate becomes $1.50 per million input and $7.50 per million output — exactly double.

Even after the doubling it stays cheaper than Claude Opus 5 at $5.00 / $25.00 and GPT-5.6 Sol at $4.00 / $20.00. But if you are building a cost model for next year, use the January numbers, not today's.


Advertisement

Gemini 3.8 Flash vs Gemini 3.7 Flash

Same price, three weeks apart, and the gains are real rather than cosmetic:

  • DeepSWE v1.1: 65.3% → 73.7% (+8.4)
  • Terminal-bench 4.0: 11.2% → 19.1% (+7.9)
  • BioMysteryBench human-difficult: 43.5% → 56.5% (+13.0)
  • OSWorld-2.0: 50.6% → 59.0% (+8.4)
  • HLE-Verified: 53.6% → 54.9% (+1.3)

The biology jump of thirteen points is the largest, and 3.8 Flash is ahead on all fourteen capability rows of this table at the same price. That makes it the stronger capability choice; 3.7 Flash may still suit workloads where token usage, latency or migration risk matters more than peak scores, so measure your own before moving.


Security: Prompt Injection and Flash Cyber

Two separate releases sit under this heading: a hardening result for the general model, and a distinct model built for defensive security work.

Indirect prompt injection

Google published a Gray Swan IPI result, which measures attack success rate within k attempts — lower is better.

At ASR@15, Gemini 3.8 Flash sits at 5.5%, and Gemini 3.8 Flash Cyber at 6.0%. For comparison, Gemini 3.7 Flash was at 9.2%, Claude Opus 5 leads at 4.8%, and the weakest model on the chart, DeepSeek V4 Pro, sits at 60.1%.

Going from 9.2% to 5.5% is a 40% reduction in attack success rate, not a halving. This is an indirect prompt-injection evaluation specifically — untrusted content reaching the model through a page or document — so it says nothing about jailbreak resistance in general. Still the most useful number here if you are shipping an agent that reads the open web.

Gray Swan indirect prompt injection attack success rate, lower is better: Gemini 3.8 Flash 5.5%

Gray Swan IPI, attack success rate within 15 attempts. Lower is better. Gemini 3.8 Flash at 5.5% against 9.2% for 3.7 Flash; Claude Opus 5 leads at 4.8%.


What Gemini 3.8 Flash Cyber is

A separate model for defensive security: finding vulnerabilities and writing patches. It runs with a more permissive set of mitigations than the general model, because a model that refuses to discuss exploit code is useless to a security team.

CyberGym Pass@1, vulnerability discovery in C/C++: 86.2%. That beats Gemini 3.5 Flash Cyber at 77.5%, GPT-5.5-Cyber at 85.6%, Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

CyberGym Pass@1 vulnerability discovery in C and C++: Gemini 3.8 Flash Cyber 86.2%

CyberGym Pass@1. Flash Cyber's 86.2% is 8.7 points above the previous Cyber model and 0.6 above the nearest competitor.

Google's internal real-world vulnerability discovery evaluation, spanning 20 programming languages: 71.0%, against Gemini 3.7 Flash at 58.9% and Gemini 3.5 Flash Cyber at 46.6% — a jump of over twenty-four points in two versions. This is one internal dataset, not a general claim that it finds 71% of vulnerabilities in the wild.

Real-world vulnerability discovery across 20 languages: 71.0% for Gemini 3.8 Flash Cyber

Google's internal real-world evaluation across 20 languages: 46.6% to 71.0% across two model generations.

CWE-Bench Pass@1: 47.2%, just behind the leading frontier model at 47.8%. The interesting part is the cost axis: in Google's chart, accuracy is plotted against average cost per rollout, and Gemini 3.8 Flash Cyber reaches roughly the same accuracy as Fable 5 at about a third of the cost per rollout.

CWE-Bench Pass@1 plotted against average cost per rollout, showing Gemini 3.8 Flash Cyber on the pareto curve

CWE-Bench accuracy against cost per rollout. Flash Cyber matches Fable 5's accuracy at roughly a third of the cost — the reason this chart is worth more than the raw score.

The catch is access. Flash Cyber is not on the public API. It is released through the Fairwind Program, limited to trusted government authorities, critical infrastructure operators, and software maintainers. For most readers this is a model to know about, not one to plan around.


Advertisement

Where You Can Use Gemini 3.8 Flash Today

  • Google Antigravity
  • Gemini API — via Google AI Studio and Android Studio
  • Gemini Enterprise
  • Gemini app — Pro and Ultra subscribers
  • AI Mode in Google Search
  • Google Sheets

Gemini 3.8 Flash Cyber is available only through the Fairwind Program.


Should You Switch?

Three honest cases.

Switch, if you are on Gemini 3.7 Flash and want capability. Same price, and ahead on every row of this comparison.

Switch, if your workload is document reasoning, financial or legal analysis, chart interpretation, long video, or structured coding, and you are currently paying Opus 5 prices. You get near-parity on DeepSWE and better numbers on the professional benchmarks for 15% of the input cost.

Do not switch, if your product depends on long-horizon autonomous agents or computer use. Terminal-bench 4.0 at 19.1% against 51.8%, and OSWorld-2.0 at 59.0% against 75.4%, are gaps that pricing does not close. Run your own evaluation on your own tasks before moving anything.


Frequently Asked Questions

Which benchmark has Gemini 3.8 Flash's highest score in this comparison? Terminal-bench 2.1 at 89.4%. Scores are not comparable across benchmarks, so that is the highest number here rather than the best result. Its largest relative lead is Harvey's Legal Agent Benchmark, about 49% above Opus 5; its largest absolute lead is BioMysteryBench human-difficult, 7.1 points clear.

Is Gemini 3.8 Flash better than Claude Opus 5? It scores higher on eight of the fourteen capability rows in this comparison, at 15% of the input price. Opus 5 leads on Terminal-bench 4.0, OSWorld-2.0, GDPVal-AA v2, BioMysteryBench human-solvable and DeepSWE — and the first two are not close. Different benchmarks measure different things, so this is not an overall ranking.

How much does Gemini 3.8 Flash cost? $0.75 per million input tokens and $3.75 per million output tokens, uncached, until 31 December 2026, then $1.50 and $7.50 from 1 January 2027.

Is Gemini 3.8 Flash Cyber publicly available? No. It is distributed through the Fairwind Program to trusted government authorities, critical infrastructure operators and software maintainers.

What is HLE-Verified? A multidisciplinary expert reasoning benchmark spanning STEM and the humanities. Gemini 3.8 Flash scores 54.9%, the highest of the six models compared.


Advertisement

Where These Numbers Come From

Every score above is transcribed from Google's published evaluation charts, not from the summary text. Here is the original comparison table for reference:

Google's full Gemini 3.8 Flash comparison table across 15 benchmarks and six models

Google's launch comparison table, the source for the figures in this article.

The announcement is on Google's blog, and the per-benchmark methodology is at the deepmind.google evaluation pages linked beneath each of Google's charts.


Recap

Gemini 3.8 Flash leads eight of the fourteen capability rows in Google's launch comparison and does it at $0.75 per million input tokens, 15% of Claude Opus 5's rate. The strongest results are in professional domains — finance, law, biology — plus chart reasoning and long video, and its Gray Swan indirect-prompt-injection attack-success rate fell about 40%, from 9.2% to 5.5%.

The losses are concentrated and worth respecting. Long-horizon agent work on Terminal-bench 4.0 and computer use on OSWorld-2.0 both sit far behind Opus 5, and GPT-5.6 Sol still leads on expert PDF comprehension. Remember too that today's price doubles on 1 January 2027.

Gemini 3.8 Flash Cyber is the more unusual release: 86.2% on CyberGym and 71.0% on real-world vulnerability discovery across twenty languages, but locked behind the Fairwind Program, so for most teams it sets expectations rather than plans.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments