← Back to Comparisons
Comparison · James Park ·

Grok 4.6 vs DeepSeek V4 Pro 0813 vs Qwen3.8-Max — The Aug 12, 2026 Frontier Release Day, Compared

Grok 4.6 vs DeepSeek V4 Pro 0813 vs Qwen3.8-Max — The Aug 12, 2026 Frontier Release Day, Compared

Grok 4.6 vs DeepSeek V4 Pro 0813 vs Qwen3.8-Max — The Aug 12, 2026 Frontier Release Day, Compared

Quick Verdict

On August 12, 2026, three labs shipped frontier models within hours of each other: xAI’s Grok 4.6 (agentic coding focus, AA Intelligence Index 61), DeepSeek’s V4 Pro 0813 (GA flagship, 62.5 benchmark geometric mean, $0.435/$0.87 per 1M), and Qwen’s Qwen3.8-2.4T-A95B open weights (first Max-class open release, 2.4T params / 95B active). The timing was not an accident — HN commenters immediately flagged it: “Just after DeepSeek-V4-Pro-0813 published, is this on purpose?” / “I think both are after Qwen 3.8 Max release.”

Pick by workload, not by benchmark:

  • Grok 4.6 — fastest to a working product, cheapest frontier on Cursor subscription (2x included usage first week), best for interactive/visual builds. $2/$6 per 1M.
  • DeepSeek V4 Pro 0813 — the value king for high-volume agentic loops: 60x effective discount via cache reads, open weights, 1M context. $0.435/$0.87 per 1M.
  • Qwen3.8-Max (2.4T-A95B) — the open-weights flagship: PaperBench 93.0 and Terminal Bench 86.6 beat Opus 4.8, but ~5TB BF16 makes it an infrastructure project. API from $2/$6 per 1M.

The Release Day Timeline

  • Qwen3.8-2.4T-A95B weights → Hugging Face (Aug 12, ~10:24 UTC), 443 HN points — first Max-class open release
  • DeepSeek V4 Pro 0813 → OpenRouter + official API (Aug 12), 671 HN points — GA stamp on V4 Pro
  • Grok 4.6 → Cursor, Grok Build, API (Aug 12), 349 + 295 HN points across two threads — the “Fable-class, 2x usage” launch

The consensus reading in the community: everyone timed around Qwen’s weight drop, and DeepSeek and xAI both accelerated to share the news cycle.

Benchmark Comparison

Artificial Analysis Intelligence Index (composite of 9 benchmarks)

ModelAA Intelligence Index
Fable 5 Max62
Grok 4.661
GPT-5.6 Sol Max61
Grok 4.556

Grok 4.6 matches GPT-5.6 Sol on the composite index — xAI’s claim of “frontier intelligence” holds up at the aggregate level.

Agentic coding benchmarks (xAI published table)

BenchmarkGrok 4.6Grok 4.5GPT-5.6 Sol MaxFable 5 Max
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Grok 4.6 beats GPT-5.6 Sol on 6 of 9 rows and trails Fable 5 narrowly overall — a genuine Fable-adjacent agentic coder, with the caveat that Terminal-Bench v3.0 (26%) is a real weak spot.

DeepSeek V4 Pro 0813 vs the field (community-collected geometric mean)

ModelGeometric Mean (7 benchmarks)
GPT-5.6 Sol65.5
Fable 5 (w/ fallback)64.5
Opus 564.0
DS-V4-Pro 081362.5
Kimi-K362.3
DS-V4-Flash 073155.8
GLM-5.247.3

Qwen3.8-Max standout scores

  • PaperBench: 93.0 (beats Opus 4.8’s 80.3, Fable 5’s 88.8, Sol’s 90.5)
  • Terminal Bench 2.1: 86.6 (beats Opus 4.8’s 84.6 and Fable 5’s 84.6)
  • GPQA Diamond: 92.6 (frontier tier)
  • Weak spots: DeepSWE 56.6, HLE w/ tools 56.2

Pricing Comparison (per 1M tokens)

ModelInputOutputCached InputEffective Agentic Cost
DeepSeek V4 Pro 0813$0.435$0.87$0.003625~$0.000875/request with cache hits
Grok 4.6$2$6— (fast variant 2x)Cheap via Cursor subscription (2x included usage, first week)
Qwen3.8-Max (API)$2$6$2/$6 via Qwen Cloud
GPT-5.6 Sol (ref)~$12.50~$60Premium
Fable 5 (ref)~$15~$75Premium

Grok 4.6’s pricing starts at $2/$6 per 1M — 10x cheaper than the US frontier tier and explicitly positioned against Kimi-K3’s API price. DeepSeek remains 4.6x cheaper than Grok on list price, and the cache-read discount widens the gap dramatically for agent loops.

Ecosystem and Fit

Grok 4.6 — the Cursor-native pick

  • Available in Cursor and Grok Build day one, with 2x included usage for the first week
  • API via x.ai, OpenRouter, Vercel, Cloudflare
  • Trained for “long-running agents and ambitious interactive/visual work”; xAI says it produces “stronger first passes on visual and interactive projects” — useful for turning a product idea into a working first version in one pass
  • Community test (Codex CLI, same feature): Grok 4.6 finished in 3m 18s, $1.41, no bug — vs DeepSeek’s 12m 02s, $0.12, with a bug
  • Watch out: Terminal-Bench v3.0 at 26% is the weakest agentic number in the class; the xAI-Musk association is a hard pass for some users (“Won’t touch xAI things”)

DeepSeek V4 Pro 0813 — the value pick for volume

  • Open weights, 1M context, no vision
  • OpenAI-compatible API everywhere (OpenRouter direct-hosted)
  • Real-world users pair it with planning models: “I will typically plan using Opus or GLM, then implement with DSF” — 0813 upgrades that execution tier
  • Enterprise adoption barrier: data policy (trains on prompts) and geopolitical risk
  • Watch out: no vision; single-benchmark wins (Cybergym 83.3, AutomationBench 31.8) but below Fable on DeepSWE/Toolathlon

Qwen3.8-Max / 2.4T-A95B — the open-weights flagship

  • Only one of the three where you can download the frontier (BF16 ~5TB, FP8 ~2.5TB, 1-bit quant 397GB)
  • Hybrid Gated DeltaNet architecture — the most interesting research artifact of the day
  • API version adds vision, 1M context, tools; open card is text/thinking only
  • Qwen3.8-27B (the local-friendly size) announced for Friday
  • Watch out: ~5TB serving cost, non-MIT license, “hobbled” open card

Three-Way Verdict by Use Case

Use CaseWinnerWhy
Cursor/IDE coding agentGrok 4.6Native integration, 2x usage week one, beats Sol on 6/9 agentic benchmarks
High-volume agent APIsDeepSeek V4 Pro 081360x cache-read economics, open weights, 1M context
Self-hosting / researchQwen3.8-2.4T-A95BOnly open frontier weights; DeltaNet architecture to study
Enterprise (regulated)Grok 4.6 (US) / Qwen-Max API (CN)DeepSeek’s training-on-prompts policy is the blocker
Local consumer hardwareNone yetQwen3.8-27B Friday; DeepSeek V4 Flash 0731 remains the local-value king

Community Sentiment Snapshot

  • On the timing: “Just after DeepSeek-V4-Pro-0813 published, is this on purpose?” — HN
  • On Grok 4.6: “Fable level performance, faster and significantly cheaper. Wow!” / “Benchmarks are almost useless. Having said that, Grok 4.6 (1.5T params) is without a doubt way smaller than Fable” / “Nazi model looks really good on the benchmarks”
  • On DeepSeek: “Competitive with opus 4.8 but weaker than sol or fable. About 20x cheaper.” / “Deepseek seems to have gotten too cheap… my credits balance barely moves”
  • On Qwen: “The card looks almost too good to be true” / “Not seeing the upside versus K3 here, especially with the intentional capability loss. Read the room, Qwen.”
  • On the meta-trend: “Within 2 months of Fable releasing, all the major labs suddenly had Fable-level models” — with competing explanations from distillation, compute buildouts, and “benchmaxxing” skepticism

FAQ

Which model is best for coding in Cursor? Grok 4.6 is the native pick (in Cursor day one, 2x usage for a week), and it beats GPT-5.6 Sol on 6 of 9 agentic benchmarks in xAI’s published table. DeepSeek V4 Pro 0813 via OpenRouter is the budget alternative.

Is DeepSeek V4 Pro 0813 competitive with Grok 4.6? On aggregate benchmarks yes — 62.5 geometric mean vs Grok’s Fable-adjacent index score — at 1/4.6 the list price and ~60x effective discount with cache reads. But Grok beat it decisively in one community Codex CLI test (3m18s/$1.41/no bug vs 12m02s/$0.12/bug).

Can I run Qwen3.8-2.4T-A95B at home? Not practically — ~5TB BF16, ~2.5TB FP8, 397GB at 1-bit. Wait for Qwen3.8-27B (Friday) for consumer hardware.

Why did all three release the same day? The community consensus: Qwen’s weight drop was the anchor, and DeepSeek and xAI accelerated to share the news cycle. It’s the new “release day” pattern for frontier labs.

Which has the best long-context story? DeepSeek (1M native) and Qwen3.8-Max API (1M default) tie; the open Qwen card is 262K native extendable to ~1M. Grok 4.6’s context specs weren’t a launch talking point.

Are these benchmarks trustworthy? Take the vendor tables with salt. HN’s recurring critique: “they are benchmaxxing” — RL-tuned to specific evals. The geometric-mean community analysis and real-world Codex CLI tests are more reliable signals than any single headline number.