Two frontier models shipped inside 24 hours of each other. xAI released Grok 4.6 on August 12, 2026, aimed at long-running agents and interactive work. Google released Gemini 3.7 Flash on August 13, 2026, calling it their "most intelligent workhorse model yet for coding and agents."
Both aim at the same job. They run agents that stay on task for many steps, and they write code you can ship. So which one should you reach for? The honest answer takes some care, because the two launches barely measure the same things.
In this blog, we will put both releases side by side, find the one benchmark they genuinely share, and separate the real signal from the marketing.
What Grok 4.6 Actually Claims
Grok 4.6 builds on Grok 4.5. The focus is long-running agents and richer visual work. xAI says it stays with hard tasks for many steps, whether that means researching a topic, working across a codebase, or turning an idea into a finished app.
The headline claim is that it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks.

Grok 4.6 scores 61, tying GPT-5.6 Sol Max and sitting one point behind Fable 5 Max at 62. Grok 4.5 scored 56, so the jump generation over generation is real.
Where Grok 4.6 genuinely leads is knowledge work rather than raw coding.

GDPVal-AA v2 scores real work tasks with an Elo rating. Grok 4.6 takes first place at 1753, ahead of Fable 5 Max at 1741 and GPT-5.6 Sol Max at 1728. It also leads AA-Briefcase at 1577. On Harvey LAB it scores 15.8%, far ahead of GPT-5.6 Sol Max at 2.5%.
Here is the full published table.

Read that table closely and a pattern shows up. Grok 4.6 wins on knowledge and document work. On pure coding it keeps up, but does not lead. DeepSWE v1.1 at 65.9% sits behind GPT-5.6 Sol Max at 73%. Terminal-Bench v3.0 at 26% sits well behind the same model at 34.6%.

Coding agent benchmarks land in the middle of the pack.

CursorBench v3.2 puts Grok 4.6 at 69.9%, just under Fable 5 Max at 70.5%. FrontierCode v1.1 Extended shows 61.3% against Fable's 63.6%.

On training, xAI ran a longer extra pass than Grok 4.5. They used curated model-made data for reasoning, a better optimizer, and agent RL across kernel work, web development, and CAD. They also saw more self-testing on long runs, with the model checking its own work before moving on.
What Gemini 3.7 Flash Actually Claims
Google frames this one differently, and that matters. Gemini 3.7 Flash is a workhorse model, not a flagship. It landed just three weeks after Gemini 3.6 Flash. Almost every number Google published compares it to its own last model, not to rivals.

The generational gains are large. On DeepSWE v1.1, 3.7 Flash scores 65.3% against 49.0% for 3.6 Flash. On FrontierCode 1.1 Main, it reaches 43.6% against 34.4%.

Web development improved too. On Arena.ai's WebDev Arena, 3.7 Flash posts an Elo of 1588 against 1538 for 3.6 Flash, and Google says it produces more feature-complete apps in fewer prompts.

For document-heavy fields, the jumps are the most dramatic in the release.

On GDP.pdf, a test of complex document comprehension, 3.7 Flash hits 34.0% against 22.0%. On AutomationBench, which measures real business workflow completion, it reaches 30.4% against 17.0%, close to doubling.

Here is the full benchmark set Google published.

The Four Benchmarks They Both Report
This is where most comparisons go wrong, so let's be careful.
The two announcements look like they measure different things, and mostly they do. But once you read Google's full benchmark table, four rows line up exactly with rows in xAI's table: same benchmark, same version.
| Benchmark | Grok 4.6 | Gemini 3.7 Flash | Winner |
|---|---|---|---|
| AA Intelligence Index | 61 | 56 | Grok 4.6 |
| GDPVal-AA v2 (Elo) | 1753 | 1525 | Grok 4.6 |
| DeepSWE v1.1 | 65.9% | 65.3% | Effectively tied |
| Terminal-bench 3.0 | 26.0% | 14.9% | Grok 4.6 |
Grok 4.6 wins three of the four, and two of those wins are wide. On GDPVal-AA v2 it leads by 228 Elo points, which is not a rounding difference. On Terminal-bench 3.0 it nearly doubles Gemini's score.
But look closely at the third row. On DeepSWE v1.1, the long-horizon software engineering test, the gap is 0.6 percentage points. That is a tie. A flagship-priced model and a workhorse model land on the same agentic coding score.
That single row is the one most people should care about, because it is the closest thing here to "can this model work through a real codebase task." Hold it in mind, because the pricing section is about to make it matter.
One caveat worth stating. Even a shared benchmark can be run differently. xAI notes its DeepSWE figure came from a mini-swe-agent harness run by Datacurve, and Google does not specify a matching harness in its table. Treat a 0.6-point gap as a tie rather than a win for anyone.
The FrontierCode Trap
You will see posts comparing Grok 4.6's 61.3% on FrontierCode against Gemini 3.7 Flash's 43.6% and concluding Grok is far ahead on production code. That comparison is wrong.
Look at the exact labels:
- Grok 4.6 reports FrontierCode v1.1 (Extended) at 61.3%
- Gemini 3.7 Flash reports FrontierCode 1.1 Main at 43.6%
Extended and Main are different slices of the benchmark, and they are not equally hard. A score on one tells you nothing about a score on the other. Neither company published the missing variant. So on FrontierCode, these two models cannot be compared at all.
The same trap appears with Harvey. Grok 4.6 reports Harvey LAB (Vals) at 15.8%. Gemini 3.7 Flash reports Harvey LAB-AA at 90.7%. Those scores are 75 points apart, which alone should tell you they are not the same test.
Here, we can see why reading benchmark labels carefully matters more than reading the numbers. Same benchmark family, same version number, completely different test.
Pricing: Where the Real Gap Is
Now the part that changes decisions.
| Grok 4.6 | Gemini 3.7 Flash | |
|---|---|---|
| Input, per 1M tokens | 0.75 (introductory) | |
| Output, per 1M tokens | 3.75 (introductory) | |
| Price after promo | Unchanged | 7.50 out from Jan 1, 2027 |
| Fast variant | 2x the price | Not offered |
At introductory pricing, Gemini 3.7 Flash costs about 63% less on input and 38% less on output than Grok 4.6, for a DeepSWE score within 0.6 points.
That is the single most important line in this article. On the one shared benchmark where the two models tie, DeepSWE, you are paying a large premium for Grok 4.6 to land in the same place. On the three where Grok wins, the premium buys you something real.

Two things keep this fair. First, Gemini's price is an intro rate that ends on December 31, 2026. From January 1, 2027 it moves to 7.50 output. At that point output tokens cost more than Grok 4.6. If your work is output-heavy and long-lived, plan for that now.
Second, one benchmark is one benchmark. Grok 4.6's wins on GDPVal-AA, AA-Briefcase, and Harvey LAB are real, and Gemini published nothing comparable on those.
Where Each One Fits
Let me tabulate this for your better understanding.
| Your workload | Better fit | Why |
|---|---|---|
| High-volume agentic coding | Gemini 3.7 Flash | Same DeepSWE score, far lower cost per token |
| Legal, finance, dense documents | Grok 4.6 | Leads GDPVal-AA, AA-Briefcase, Harvey LAB |
| Terminal and shell agents | Neither, look at GPT-5.6 | Both trail badly on Terminal-bench 3.0: 26% and 14.9% |
| Web app and UI generation | Gemini 3.7 Flash | WebDev Arena 1588, strong design adherence |
| Long visual and interactive builds | Grok 4.6 | xAI's stated focus, stronger first passes |
| Already inside Google Cloud | Gemini 3.7 Flash | AI Studio, Antigravity, Gemini Enterprise |
| Already inside Cursor | Grok 4.6 | Shipped in Cursor day one, 2x usage first week |
The uncomfortable truth in that table is Terminal-bench 3.0. Grok 4.6 scores 26%, well behind GPT-5.6 Sol Max at 34.6%. Gemini 3.7 Flash scores 14.9%, behind GPT-5.6 Terra at 20.8% in its own table. If your agent lives in a shell, neither of this week's releases is your answer.
How to Get Access
Grok 4.6 is available in Cursor, Grok Build, and the xAI API, plus OpenRouter, Vercel, and Cloudflare. xAI offered 2x included usage in Grok Build and Cursor for the first week.
Gemini 3.7 Flash is available in Google AI Studio, Google Antigravity, and Android Studio for developers; Gemini Enterprise for companies; and powers Gemini Spark for AI Pro and Ultra subscribers across 160+ countries.
The fairest test is your own workload with a fixed set of prompts. Measure cost per finished task, not tokens per second. A benchmark table tells you what a model can do. Your own eval tells you what it will do for you.
Recap
This is how Grok 4.6 and Gemini 3.7 Flash compare. Grok 4.6 is the broader model, leading on knowledge work with a top GDPVal-AA score of 1753, a first-place AA-Briefcase result, and the strongest Harvey LAB number in its table. Gemini 3.7 Flash is a workhorse that closed most of the gap to flagship coding performance in a single three-week release cycle.
They meet on four shared benchmarks. Grok 4.6 wins three of them, leading by 228 Elo on GDPVal-AA v2 and nearly doubling Gemini on Terminal-bench 3.0. But on DeepSWE v1.1, the long-horizon coding test, they tie at 65.9% against 65.3%. At the intro price, Gemini reaches that tie for about a third of the input cost, which makes it the default for high-volume agent coding until the intro rate ends on December 31, 2026.
Be careful with the FrontierCode numbers circulating online. Grok reports the Extended variant and Google reports Main, so those two scores were never comparable. And if your agents work in a terminal, look past both of these releases, because Terminal-Bench remains the weak spot.
Sources: Introducing Grok 4.6 (xAI) and Introducing Gemini 3.7 Flash (Google). All figures are the developers' own published numbers.