Grok 4.6 vs Gemini 3.7 Flash on Benchmarks and Pricing

Grok 4.6 vs Gemini 3.7 Flash compared on the four benchmarks both companies actually publish, plus pricing, the FrontierCode label trap, and which model fits which workload.

Aug 14, 2026Updated Aug 26, 202610 min readFollow

Topics You Will Master

What Grok 4.6 and Gemini 3.7 Flash each claim, straight from the release notes
The four benchmarks both companies report, and who wins each one
Why the FrontierCode numbers you will see compared online are not comparable
What each model costs, and where the price gap actually changes your decision

Two frontier models shipped inside 24 hours of each other. xAI released Grok 4.6 on August 12, 2026, aimed at long-running agents and interactive work. Google released Gemini 3.7 Flash on August 13, 2026, calling it their "most intelligent workhorse model yet for coding and agents."

Both aim at the same job. They run agents that stay on task for many steps, and they write code we can ship. So which one should we reach for? The honest answer takes some care, because the two launches barely measure the same things.

Bestseller

Fine Tuning LLM with Hugging Face Transformers for NLP

Learn transformer architecture fundamentals and fine-tune LLMs with custom datasets.

95% off$199.99
$9.99Enroll now 30 day refund, lifetime access

In this blog, we will put both releases side by side, find the one benchmark they genuinely share, and separate the real signal from the marketing.

What Does Grok 4.6 Claim?

Grok 4.6 builds on Grok 4.5. The focus is long-running agents and richer visual work. xAI says it stays with hard tasks for many steps, whether that means researching a topic, working across a codebase, or turning an idea into a finished app.

The headline claim is that it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks.

Grok 4.6 AA Intelligence Index

Grok 4.6 scores 61, tying GPT-5.6 Sol Max and sitting one point behind Fable 5 Max at 62. Grok 4.5 scored 56, so the jump generation over generation is real.

Where Grok 4.6 genuinely leads is knowledge work rather than raw coding.

Grok 4.6 GDPVal-AA Elo

GDPVal-AA v2 scores real work tasks with an Elo rating. Grok 4.6 takes first place at 1753, ahead of Fable 5 Max at 1741 and GPT-5.6 Sol Max at 1728. It also leads AA-Briefcase at 1577. On Harvey LAB it scores 15.8%, far ahead of GPT-5.6 Sol Max at 2.5%.

Here is the full published table.

Grok 4.6 full evaluation table

If we read that table closely, a pattern shows up. Grok 4.6 wins on knowledge and document work. On pure coding it keeps up, but does not lead. DeepSWE v1.1 at 65.9% sits behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 at 26% sits well behind the same model at 34.6%.

Grok 4.6 DeepSWE score

Coding agent benchmarks land in the middle of the pack.

Grok 4.6 CursorBench score

CursorBench v3.2 puts Grok 4.6 at 69.9%, just under Fable 5 Max at 70.5%. FrontierCode v1.1 Extended shows 61.3% against Fable's 63.6%.

Grok 4.6 FrontierCode Extended score

On training, xAI ran a longer extra pass than Grok 4.5. They used curated model-made data for reasoning, a better optimizer, and agent RL across kernel work, web development, and CAD. They also saw more self-testing on long runs, with the model checking its own work before moving on.

Advertisement

What Does Gemini 3.7 Flash Claim?

Google frames this one differently, and that matters. Gemini 3.7 Flash is a workhorse model, not a flagship. It landed just three weeks after Gemini 3.6 Flash. Almost every number Google published compares it to its own last model, not to rivals.

Gemini 3.7 Flash DeepSWE

The generational gains are large. On DeepSWE v1.1, 3.7 Flash scores 65.3% against 49.0% for 3.6 Flash. On FrontierCode 1.1 Main, it reaches 43.6% against 34.4%.

Gemini 3.7 Flash FrontierCode

Web development improved too. On Arena.ai's WebDev Arena, 3.7 Flash posts an Elo of 1588 against 1538 for 3.6 Flash, and Google says it produces more feature-complete apps in fewer prompts.

Gemini 3.7 Flash WebDev Arena

For document-heavy fields, the jumps are the most dramatic in the release.

Gemini 3.7 Flash GDP.pdf

On GDP.pdf, a test of hard document reading, 3.7 Flash hits 34.0% against 22.0%. On AutomationBench, which measures real business workflow completion, it reaches 30.4% against 17.0%, close to doubling.

Gemini 3.7 Flash AutomationBench

Here is the full benchmark set Google published.

Gemini 3.7 Flash detailed benchmarks

Which Benchmarks Do Both Models Report?

This is where most comparisons go wrong, so let's be careful.

The two announcements look like they measure different things, and mostly they do. But once we read Google's full benchmark table, four rows line up exactly with rows in xAI's table: same benchmark, same version.

Benchmark Grok 4.6 Gemini 3.7 Flash Winner
AA Intelligence Index 61 56 Grok 4.6
GDPVal-AA v2 (Elo) 1753 1525 Grok 4.6
DeepSWE v1.1 65.9% 65.3% Effectively tied
Terminal-bench 3.0 26.0% 14.9% Grok 4.6

Grok 4.6 wins three of the four, and two of those wins are wide. On GDPVal-AA v2 it leads by 228 Elo points, which is not a rounding difference. On Terminal-bench 3.0 it nearly doubles Gemini's score.

But let's look closely at the third row. On DeepSWE v1.1, the long-horizon software engineering test, the gap is 0.6 percentage points. That is a tie. A flagship-priced model and a workhorse model land on the same agentic coding score.

That single row is the one most people should care about, because it is the closest thing here to "can this model work through a real codebase task." We should keep it in mind, because the pricing section is about to make it matter.

One caveat worth stating. Even a shared benchmark can be run differently. xAI notes its DeepSWE figure came from a mini-swe-agent harness run by Datacurve, and Google does not specify a matching harness in its table. So we should treat a 0.6-point gap as a tie rather than a win for anyone.

Advertisement

Why Are the FrontierCode Numbers Not Comparable?

Many posts online compare Grok 4.6's 61.3% on FrontierCode against Gemini 3.7 Flash's 43.6% and conclude Grok is far ahead on production code. That comparison is wrong.

Let's look at the exact labels:

  • Grok 4.6 reports FrontierCode v1.1 (Extended) at 61.3%
  • Gemini 3.7 Flash reports FrontierCode 1.1 Main at 43.6%

Extended and Main are different slices of the benchmark, and they are not equally hard. A score on one tells us nothing about a score on the other. Neither company published the missing variant. So on FrontierCode, these two models cannot be compared at all.

The same trap appears with Harvey. Grok 4.6 reports Harvey LAB (Vals) at 15.8%. Gemini 3.7 Flash reports Harvey LAB-AA at 90.7%. Those scores are 75 points apart, which alone should tell us they are not the same test.

Here, we can see why reading benchmark labels carefully matters more than reading the numbers. Same benchmark family, same version number, completely different test.

How Much Does Each Model Cost?

Now the part that changes decisions.

Grok 4.6 Gemini 3.7 Flash
Input, per 1M tokens 0.75 (introductory)
Output, per 1M tokens 3.75 (introductory)
Price after promo Unchanged 7.50 out from Jan 1, 2027
Fast variant 2x the price Not offered

At introductory pricing, Gemini 3.7 Flash costs about 63% less on input and 38% less on output than Grok 4.6, for a DeepSWE score within 0.6 points.

That is the single most important line in this article. On the one shared benchmark where the two models tie, DeepSWE, we are paying a large premium for Grok 4.6 to land in the same place. On the three where Grok wins, the premium buys us something real.

Gemini 3.7 Flash performance to cost

Two things keep this fair. First, Gemini's price is an intro rate that ends on December 31, 2026. From January 1, 2027 it moves to 7.50 output. At that point output tokens cost more than Grok 4.6. If our work is output-heavy and long-lived, we should plan for that now.

Second, one benchmark is one benchmark. Grok 4.6's wins on GDPVal-AA, AA-Briefcase, and Harvey LAB are real, and Gemini published nothing comparable on those.

Advertisement

Which Model Fits Which Workload?

Let me tabulate this for your better understanding.

Your workload Better fit Why
High-volume agentic coding Gemini 3.7 Flash Same DeepSWE score, far lower cost per token
Legal, finance, dense documents Grok 4.6 Leads GDPVal-AA, AA-Briefcase, Harvey LAB
Terminal and shell agents Neither, look at GPT-5.6 Both trail badly on Terminal-bench 3.0: 26% and 14.9%
Web app and UI generation Gemini 3.7 Flash WebDev Arena 1588, strong design adherence
Long visual and interactive builds Grok 4.6 xAI's stated focus, stronger first passes
Already inside Google Cloud Gemini 3.7 Flash AI Studio, Antigravity, Gemini Enterprise
Already inside Cursor Grok 4.6 Shipped in Cursor day one, 2x usage first week

The uncomfortable truth in that table is Terminal-bench 3.0. Grok 4.6 scores 26%, well behind GPT-5.6 Sol Max at 34.6%. Gemini 3.7 Flash scores 14.9%, behind GPT-5.6 Terra at 20.8% in its own table. If our agent lives in a shell, neither of this week's releases is the answer.

How to Get Access

Grok 4.6 is available in Cursor, Grok Build, and the xAI API, plus OpenRouter, Vercel, and Cloudflare. xAI offered 2x included usage in Grok Build and Cursor for the first week.

Gemini 3.7 Flash is available in Google AI Studio, Google Antigravity, and Android Studio for developers; Gemini Enterprise for companies; and powers Gemini Spark for AI Pro and Ultra subscribers across 160+ countries.

The fairest test is our own workload with a fixed set of prompts. We should measure cost per finished task, not tokens per second. A benchmark table tells us what a model can do. Our own eval tells us what it will do for us.

Conclusion

We put Grok 4.6 and Gemini 3.7 Flash side by side using only the numbers both companies published. Grok 4.6 is the broader model. It leads on knowledge work with a top GDPVal-AA score of 1753, a first-place AA-Briefcase result, and the strongest Harvey LAB number in its table. Gemini 3.7 Flash is a workhorse that closed most of the gap to flagship coding performance in a single three-week release cycle.

Key takeaways:

  • The two launches share only four benchmarks. Grok 4.6 wins three of them, leading by 228 Elo on GDPVal-AA v2 and nearly doubling Gemini on Terminal-bench 3.0.
  • On DeepSWE v1.1, the long-horizon coding test, they tie at 65.9% against 65.3%. Gemini reaches that tie for about a third of the input cost.
  • That makes Gemini 3.7 Flash the default for high-volume agent coding until the intro rate ends on December 31, 2026.
  • The FrontierCode numbers circulating online are not comparable. Grok reports the Extended variant and Google reports Main, so those two scores were never the same test.
  • If our agents work in a terminal, neither release is the answer, because both trail badly on Terminal-bench 3.0.

Next steps:

This is how Grok 4.6 and Gemini 3.7 Flash compare. We started with what each release claims, we found the four benchmarks they truly share, and we finished with the pricing gap that decides which one to pick.

Sources: Introducing Grok 4.6 (xAI) and Introducing Gemini 3.7 Flash (Google). All figures are the developers' own published numbers.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments