Grok 4.6 vs Gemini 3.7 Flash: Benchmarks, Pricing, and Which One to Use

Grok 4.6 vs Gemini 3.7 Flash compared on the four benchmarks both companies actually publish, plus pricing, the FrontierCode label trap, and which model fits which workload.

Aug 14, 20269 min readFollow

Topics You Will Master

What Grok 4.6 and Gemini 3.7 Flash each claim, straight from the release notes
The four benchmarks both companies report, and who wins each one
Why the FrontierCode numbers you will see compared online are not comparable
What each model costs, and where the price gap actually changes your decision

Two frontier models shipped inside 24 hours of each other. xAI released Grok 4.6 on August 12, 2026, aimed at long-running agents and interactive work. Google released Gemini 3.7 Flash on August 13, 2026, calling it their "most intelligent workhorse model yet for coding and agents."

Both aim at the same job. They run agents that stay on task for many steps, and they write code you can ship. So which one should you reach for? The honest answer takes some care, because the two launches barely measure the same things.

Bestseller

Master LangGraph and LangChain

Agentic RAG and Chatbot, AI Agent with LangChain v1, Qwen3, Gemma3, DeepSeek-R1, LLAMA 3.2, FAISS Vector Database

Enroll on Udemy 30 day refund, lifetime access

In this blog, we will put both releases side by side, find the one benchmark they genuinely share, and separate the real signal from the marketing.


What Grok 4.6 Actually Claims

Grok 4.6 builds on Grok 4.5. The focus is long-running agents and richer visual work. xAI says it stays with hard tasks for many steps, whether that means researching a topic, working across a codebase, or turning an idea into a finished app.

The headline claim is that it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks.

Grok 4.6 AA Intelligence Index

Grok 4.6 scores 61, tying GPT-5.6 Sol Max and sitting one point behind Fable 5 Max at 62. Grok 4.5 scored 56, so the jump generation over generation is real.

Where Grok 4.6 genuinely leads is knowledge work rather than raw coding.

Grok 4.6 GDPVal-AA Elo

GDPVal-AA v2 scores real work tasks with an Elo rating. Grok 4.6 takes first place at 1753, ahead of Fable 5 Max at 1741 and GPT-5.6 Sol Max at 1728. It also leads AA-Briefcase at 1577. On Harvey LAB it scores 15.8%, far ahead of GPT-5.6 Sol Max at 2.5%.

Here is the full published table.

Grok 4.6 full evaluation table

Read that table closely and a pattern shows up. Grok 4.6 wins on knowledge and document work. On pure coding it keeps up, but does not lead. DeepSWE v1.1 at 65.9% sits behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 at 26% sits well behind the same model at 34.6%.

Grok 4.6 DeepSWE score

Coding agent benchmarks land in the middle of the pack.

Grok 4.6 CursorBench score

CursorBench v3.2 puts Grok 4.6 at 69.9%, just under Fable 5 Max at 70.5%. FrontierCode v1.1 Extended shows 61.3% against Fable's 63.6%.

Grok 4.6 FrontierCode Extended score

On training, xAI ran a longer extra pass than Grok 4.5. They used curated model-made data for reasoning, a better optimizer, and agent RL across kernel work, web development, and CAD. They also saw more self-testing on long runs, with the model checking its own work before moving on.


What Gemini 3.7 Flash Actually Claims

Google frames this one differently, and that matters. Gemini 3.7 Flash is a workhorse model, not a flagship. It landed just three weeks after Gemini 3.6 Flash. Almost every number Google published compares it to its own last model, not to rivals.

Gemini 3.7 Flash DeepSWE

The generational gains are large. On DeepSWE v1.1, 3.7 Flash scores 65.3% against 49.0% for 3.6 Flash. On FrontierCode 1.1 Main, it reaches 43.6% against 34.4%.

Gemini 3.7 Flash FrontierCode

Web development improved too. On Arena.ai's WebDev Arena, 3.7 Flash posts an Elo of 1588 against 1538 for 3.6 Flash, and Google says it produces more feature-complete apps in fewer prompts.

Gemini 3.7 Flash WebDev Arena

For document-heavy fields, the jumps are the most dramatic in the release.

Gemini 3.7 Flash GDP.pdf

On GDP.pdf, a test of complex document comprehension, 3.7 Flash hits 34.0% against 22.0%. On AutomationBench, which measures real business workflow completion, it reaches 30.4% against 17.0%, close to doubling.

Gemini 3.7 Flash AutomationBench

Here is the full benchmark set Google published.

Gemini 3.7 Flash detailed benchmarks


The Four Benchmarks They Both Report

This is where most comparisons go wrong, so let's be careful.

The two announcements look like they measure different things, and mostly they do. But once you read Google's full benchmark table, four rows line up exactly with rows in xAI's table: same benchmark, same version.

Benchmark Grok 4.6 Gemini 3.7 Flash Winner
AA Intelligence Index 61 56 Grok 4.6
GDPVal-AA v2 (Elo) 1753 1525 Grok 4.6
DeepSWE v1.1 65.9% 65.3% Effectively tied
Terminal-bench 3.0 26.0% 14.9% Grok 4.6

Grok 4.6 wins three of the four, and two of those wins are wide. On GDPVal-AA v2 it leads by 228 Elo points, which is not a rounding difference. On Terminal-bench 3.0 it nearly doubles Gemini's score.

But look closely at the third row. On DeepSWE v1.1, the long-horizon software engineering test, the gap is 0.6 percentage points. That is a tie. A flagship-priced model and a workhorse model land on the same agentic coding score.

That single row is the one most people should care about, because it is the closest thing here to "can this model work through a real codebase task." Hold it in mind, because the pricing section is about to make it matter.

One caveat worth stating. Even a shared benchmark can be run differently. xAI notes its DeepSWE figure came from a mini-swe-agent harness run by Datacurve, and Google does not specify a matching harness in its table. Treat a 0.6-point gap as a tie rather than a win for anyone.


The FrontierCode Trap

You will see posts comparing Grok 4.6's 61.3% on FrontierCode against Gemini 3.7 Flash's 43.6% and concluding Grok is far ahead on production code. That comparison is wrong.

Look at the exact labels:

  • Grok 4.6 reports FrontierCode v1.1 (Extended) at 61.3%
  • Gemini 3.7 Flash reports FrontierCode 1.1 Main at 43.6%

Extended and Main are different slices of the benchmark, and they are not equally hard. A score on one tells you nothing about a score on the other. Neither company published the missing variant. So on FrontierCode, these two models cannot be compared at all.

The same trap appears with Harvey. Grok 4.6 reports Harvey LAB (Vals) at 15.8%. Gemini 3.7 Flash reports Harvey LAB-AA at 90.7%. Those scores are 75 points apart, which alone should tell you they are not the same test.

Here, we can see why reading benchmark labels carefully matters more than reading the numbers. Same benchmark family, same version number, completely different test.


Pricing: Where the Real Gap Is

Now the part that changes decisions.

Grok 4.6 Gemini 3.7 Flash
Input, per 1M tokens 0.75 (introductory)
Output, per 1M tokens 3.75 (introductory)
Price after promo Unchanged 7.50 out from Jan 1, 2027
Fast variant 2x the price Not offered

At introductory pricing, Gemini 3.7 Flash costs about 63% less on input and 38% less on output than Grok 4.6, for a DeepSWE score within 0.6 points.

That is the single most important line in this article. On the one shared benchmark where the two models tie, DeepSWE, you are paying a large premium for Grok 4.6 to land in the same place. On the three where Grok wins, the premium buys you something real.

Gemini 3.7 Flash performance to cost

Two things keep this fair. First, Gemini's price is an intro rate that ends on December 31, 2026. From January 1, 2027 it moves to 7.50 output. At that point output tokens cost more than Grok 4.6. If your work is output-heavy and long-lived, plan for that now.

Second, one benchmark is one benchmark. Grok 4.6's wins on GDPVal-AA, AA-Briefcase, and Harvey LAB are real, and Gemini published nothing comparable on those.


Where Each One Fits

Let me tabulate this for your better understanding.

Your workload Better fit Why
High-volume agentic coding Gemini 3.7 Flash Same DeepSWE score, far lower cost per token
Legal, finance, dense documents Grok 4.6 Leads GDPVal-AA, AA-Briefcase, Harvey LAB
Terminal and shell agents Neither, look at GPT-5.6 Both trail badly on Terminal-bench 3.0: 26% and 14.9%
Web app and UI generation Gemini 3.7 Flash WebDev Arena 1588, strong design adherence
Long visual and interactive builds Grok 4.6 xAI's stated focus, stronger first passes
Already inside Google Cloud Gemini 3.7 Flash AI Studio, Antigravity, Gemini Enterprise
Already inside Cursor Grok 4.6 Shipped in Cursor day one, 2x usage first week

The uncomfortable truth in that table is Terminal-bench 3.0. Grok 4.6 scores 26%, well behind GPT-5.6 Sol Max at 34.6%. Gemini 3.7 Flash scores 14.9%, behind GPT-5.6 Terra at 20.8% in its own table. If your agent lives in a shell, neither of this week's releases is your answer.


How to Get Access

Grok 4.6 is available in Cursor, Grok Build, and the xAI API, plus OpenRouter, Vercel, and Cloudflare. xAI offered 2x included usage in Grok Build and Cursor for the first week.

Gemini 3.7 Flash is available in Google AI Studio, Google Antigravity, and Android Studio for developers; Gemini Enterprise for companies; and powers Gemini Spark for AI Pro and Ultra subscribers across 160+ countries.

The fairest test is your own workload with a fixed set of prompts. Measure cost per finished task, not tokens per second. A benchmark table tells you what a model can do. Your own eval tells you what it will do for you.


Recap

This is how Grok 4.6 and Gemini 3.7 Flash compare. Grok 4.6 is the broader model, leading on knowledge work with a top GDPVal-AA score of 1753, a first-place AA-Briefcase result, and the strongest Harvey LAB number in its table. Gemini 3.7 Flash is a workhorse that closed most of the gap to flagship coding performance in a single three-week release cycle.

They meet on four shared benchmarks. Grok 4.6 wins three of them, leading by 228 Elo on GDPVal-AA v2 and nearly doubling Gemini on Terminal-bench 3.0. But on DeepSWE v1.1, the long-horizon coding test, they tie at 65.9% against 65.3%. At the intro price, Gemini reaches that tie for about a third of the input cost, which makes it the default for high-volume agent coding until the intro rate ends on December 31, 2026.

Be careful with the FrontierCode numbers circulating online. Grok reports the Extended variant and Google reports Main, so those two scores were never comparable. And if your agents work in a terminal, look past both of these releases, because Terminal-Bench remains the weak spot.

Sources: Introducing Grok 4.6 (xAI) and Introducing Gemini 3.7 Flash (Google). All figures are the developers' own published numbers.

Found this useful? Keep building with me.

New tutorials every week on YouTube: or go deeper with a full structured course.

Find this tutorial useful?

Subscribe to our YouTube channels for more practical production walk-throughs.

Discussion & Comments