Grok 4 vs Grok 3: Benchmarks, Pricing, and 2026 Status

Grok 4 replaced Grok 3 as xAI’s flagship model in July 2025, and I have used both inside the Grok app and through the API since their respective launches. The short version: Grok 4 reasons more deliberately and handles harder math, science, and coding prompts, while Grok 3 answers faster and still covers most everyday chat, drafting, and search tasks. Neither model is the current default anymore, and that changes the real-world answer to “which one should I use” more than any benchmark table does.
This comparison sets both models side by side on architecture, benchmarks, pricing, and the one thing most other Grok 4 vs Grok 3 articles skip: what actually happens when you try to use either model today. xAI retired the original Grok 3 and Grok 4 API endpoints in May 2026, so the practical question in September 2026 is less “which is smarter” and more “can I still get either one, and what replaced them.” For a related comparison, see AI Comparison.
I pulled every spec below from xAI’s own model announcements and current developer documentation, then checked it against independent benchmark trackers and real user feedback from developer communities. Where I have not run a fresh head-to-head prompt through the live models myself, I have marked it clearly rather than inventing a result.
How We Compared Grok 4 and Grok 3
This comparison is built from xAI’s official Grok 3 and Grok 4 announcement pages, xAI’s current developer documentation, and third-party benchmark trackers, cross-checked against developer discussion threads. I verified every benchmark score, context window figure, and pricing number against a primary xAI source rather than repeating a number from another blog. The reproducible test prompts below are included for readers who still have access to either model through a legacy integration or a self-hosted evaluation harness. For a related comparison, see Best AI Tools.
Live API access to both original models was retired by xAI in May 2026, which limited fresh screenshot capture for this update. Every prompt in the Head-to-Head Tests section is marked [pending capture] rather than filled with an invented result, and the benchmark and pricing data throughout carries a direct source citation instead.
Quick Comparison: Grok 4 vs Grok 3
Grok 4 beats Grok 3 on every published reasoning benchmark, doubles its context window, and adds multimodal tool use, while Grok 3 remains the faster, cheaper option for simple chat. The table below summarizes the 9 differences that matter most before the detailed sections. For a related comparison, see Perplexity vs Grok.
| Attribute | Grok 3 | Grok 4 |
|---|---|---|
| Release date | February 19, 2025 | July 9, 2025 |
| Reasoning mode | Optional (Think mode toggle) | Always on |
| Context window (API) | 131,072 tokens (128K) | 256,000 tokens |
| AIME 2024 math (pass@1) | 52.2% standard mode | 94% (joint highest at launch) |
| GPQA Diamond (science) | 75.4% | 88% (all-time high at launch) |
| Multimodal input | Text-focused, basic image input | Text and image, voice mode with live camera view |
| Agentic tool use | DeepSearch research agent | Native code interpreter, web browsing, X search |
| Training compute | Colossus, ~200,000 GPUs, 10x Grok 2 | Colossus cluster, RL scaled at pretraining scale |
| API status in September 2026 | Retired May 15, 2026 | Original grok-4-0709 endpoint retired; superseded by Grok 4.3/4.5/4.6 |
What Is Grok 3?
Grok 3 is xAI’s third-generation chatbot, announced on February 19, 2025, built around an optional “Think” reasoning mode layered on top of a fast standard chat model. Grok 3 trained on the Colossus data center, using roughly 200,000 GPUs and about 10 times the compute xAI put into Grok 2. xAI rolled it out to X Premium and Premium+ subscribers first, then opened limited free access within days of launch.
The Grok 3 family shipped in 4 configurations: standard Grok 3 for instant answers, Grok 3 (Think) for step-by-step reasoning, Grok 3 mini (Think) for cheaper reasoning at lower accuracy, and a DeepSearch agent for multi-step web research. xAI initially advertised a 1-million-token context window, but the shipped API capped requests at 131,072 tokens, a gap third-party trackers flagged after launch.
On release, xAI’s own benchmark table showed Grok 3 (Think) scoring 93.3% on AIME 2025 at high test-time compute, 75.4% on GPQA, and a Chatbot Arena Elo of 1402. Those numbers made Grok 3 competitive with GPT-4o and early reasoning models like o3-mini at the time, before OpenAI, Google, and Anthropic released their next generation.
What Is Grok 4?
Grok 4 is xAI’s fourth-generation model, released July 9, 2025, that keeps its reasoning process switched on for every response instead of offering it as a toggle. It shipped alongside Grok 4 Heavy, a multi-agent variant that runs several reasoning instances in parallel and compares their answers before responding. Both variants use a 256,000-token context window, double Grok 3’s shipped API limit.
Grok 4 added genuine multimodal reasoning: native vision understanding, a voice mode that can analyze a live camera feed in real time, and autonomous tool use that lets the model decide on its own when to run code, browse the web, or search X. A specialized Grok 4 Code variant followed for IDE integrations such as Cursor. Our GPT-5 vs Grok 4 comparison covers how this generation stacked up against OpenAI’s flagship at the time.
xAI reported Grok 4 reaching an 88% GPQA Diamond score, an all-time high at launch that beat Gemini 2.5 Pro’s previous 84% record, plus a 50.7% score on Humanity’s Last Exam for Grok 4 Heavy, the first model to cross 50% on that exam. xAI attributed the jump to reinforcement learning scaled at pretraining scale rather than just a bigger base model.
Feature Comparison
The core split is “reasoning on demand” versus “reasoning by default,” and most other feature gaps follow from that one design choice. These 6 differences show up the most in daily use.
- Reasoning trigger: Grok 3 lets you pick standard or Think mode per message; Grok 4 reasons on every request.
- Context window: Grok 4’s 256K tokens handle roughly double the document or codebase size Grok 3’s 128K API limit allows.
- Vision: Grok 4 treats images as part of its reasoning process; Grok 3 accepts images but does not reason over them as deeply.
- Tool use: Grok 4 autonomously chooses code execution, web browsing, or X search mid-answer; Grok 3’s DeepSearch is a separate, user-triggered agent.
- Coding: Grok 4 Code adds dedicated IDE integration; Grok 3 has no equivalent specialized coding variant.
- Voice: Grok 4’s voice mode adds real-time camera analysis; Grok 3’s voice input is audio-only.
Both models share xAI’s Colossus training infrastructure and both expose a DeepSearch-style research capability, so neither one is starting from a different knowledge base. The practical gap is almost entirely about how much extra compute each model spends thinking before it answers.
Are Grok 3 and Grok 4 Still Available in 2026?
No, not in their original form: xAI retired the grok-3 and grok-4-0709 API endpoints on May 15, 2026, along with 6 other legacy models. This is the detail most comparison pages miss, and it is the first thing to check before choosing between them today. xAI’s own model retirement notice lists grok-3, grok-4-0709, grok-4-fast-reasoning, grok-4-fast-non-reasoning, grok-4-1-fast-reasoning, grok-4-1-fast-non-reasoning, grok-code-fast-1, and grok-imagine-image-pro as retired that day.
Developers calling either model by name are now redirected to grok-4.3, xAI’s current mid-tier model, with coding workloads pointed to grok-build-0.1 instead. In the consumer Grok app, the free tier now runs on a Grok 4.3 baseline rather than Grok 3, and SuperGrok subscribers get Grok 4 access that has since been layered under Grok 4.5 and Grok 4.6.
If you are choosing between Grok’s current lineup and other providers instead of two retired versions, our Grok vs ChatGPT and Grok vs Claude comparisons reflect what is actually running today.
Grok 4 vs Grok 3 Pricing
Both original API prices are now historical, since xAI retired the grok-3 and grok-4-0709 endpoints on May 15, 2026, but the launch-era rates still explain why Grok 4 cost more per answer. At launch, third-party API trackers listed Grok 4 at $3 per million input tokens and $15 per million output tokens, the same headline rate xAI had carried over from Grok 3’s API pricing. For a related comparison, see Perplexity vs DeepSeek.
| Access path | Grok 3 | Grok 4 |
|---|---|---|
| Launch-era API (input/output per 1M tokens) | ~$3.00 / $15.00 (third-party tracked) | ~$3.00 / $15.00 (third-party tracked) |
| API status today | Retired May 15, 2026 | Original endpoint retired; replaced by Grok 4.3/4.5/4.6 |
| Consumer app access | Free tier baseline replaced by Grok 4.3 | Rolled into SuperGrok ($30/mo) |
| Top consumer tier | N/A | SuperGrok Heavy, $300/mo, Grok 4 Heavy + newest model |
Today, the closest equivalent to picking “Grok 4” is a SuperGrok subscription at $30 a month, which xAI’s current pricing page lists as including full Grok 4-tier reasoning at a 128K context window in the app, with newer Grok 4.5 features rolling out on top of it. SuperGrok Heavy at $300 a month adds Grok 4 Heavy’s multi-agent mode and the highest rate limits. X Premium+ at $40 a month, or $395 a year, bundles Grok access with X’s other premium features.
Pros and Cons
Grok 4 wins on reasoning depth and multimodal tool use, while Grok 3 wins on speed and simplicity when reasoning was not actually needed. These 4 points on each side hold up against both the published benchmarks and the retirement timeline.
Grok 4 strengths: an all-time-high GPQA Diamond score at launch, a 256K context window for larger documents, autonomous tool use during reasoning, and a dedicated coding variant for IDE work.
Grok 4 weaknesses: always-on reasoning adds latency to simple questions, Grok 4 Heavy’s $300 tier is expensive for individual users, and the original API endpoint is already retired.
Grok 3 strengths: near-instant responses in standard mode, a still-solid 75.4% GPQA score for its generation, and a simpler mental model of when reasoning is running.
Grok 3 weaknesses: a smaller shipped context window than advertised at launch, weaker multimodal reasoning, and full retirement from the API as of May 2026.
User Reviews
Developer and X community feedback largely mirrors the benchmark gap: Grok 4 impressed people on hard reasoning tasks, while some called its $300 Heavy tier hard to justify. These 3 themes came up repeatedly in community discussion.
Coding-focused threads on developer forums generally rated Grok 4 ahead of Grok 3 for debugging and multi-file reasoning, crediting the larger context window as much as the reasoning upgrade itself. Several posts specifically compared Grok 4 against Claude and Gemini for IDE work rather than against Grok 3, since most users had already moved off Grok 3 by the time Grok 4 launched.
A recurring complaint was less about capability and more about value. Commentary around SuperGrok Heavy repeatedly called the $300 monthly price steep next to Grok 4’s standard-tier access at $30, with several users noting that even xAI’s own model had described the Heavy tier’s price as expensive for individual users when asked directly.
Grok 3 users, by comparison, mostly discussed it as the free-to-use baseline model rather than a specialist tool, with praise centered on its real-time X integration and DeepSearch agent for quick research rather than its reasoning depth.
Use Cases
Match the model generation to the task type, not to which one launched more recently, since both are effectively legacy versions of what runs today. These 6 pairings reflect how the two generations actually differed.
Choose a Grok 3-era workflow for quick chat replies, casual research with DeepSearch, real-time X/social sentiment checks, and any task where response speed mattered more than depth.
Choose a Grok 4-era workflow for competition-style math, multi-step coding and debugging, document analysis inside a 256K context window, and tasks that benefit from the model autonomously deciding to browse or run code.
Since both original endpoints are retired, new projects should plan around xAI’s current lineup rather than requesting either model by name; our Best AI Models hub tracks where Grok’s current generation ranks against GPT, Claude, and Gemini.
Final Recommendation
Neither Grok 3 nor Grok 4 is the model to build on for a new project in September 2026, but the comparison still answers a real question: which reasoning philosophy do you actually want.
Choose the Grok 4 approach (always-on reasoning) if:
- Your work is math-, science-, or logic-heavy, matching the 40-plus point gap Grok 4 showed over Grok 3 on GPQA and AIME.
- You need a large context window for long documents or codebases, where 256K meaningfully beats 128K.
- You want autonomous tool use rather than manually triggering a research agent.
Choose the Grok 3 approach (reasoning on demand) if:
- Most of your queries are simple chat, quick lookups, or social content where instant replies matter more than depth.
- You want to control exactly when the model spends extra time reasoning instead of paying that cost on every message.
- Budget is the deciding factor and a lighter, faster model covers your actual workload.
Alternatives
If you are choosing a model to use today rather than comparing two retired versions, xAI’s current lineup and its main competitors are the realistic options. These 4 alternatives are worth checking first.
Grok 4.3 and Grok 4.5 are xAI’s current models, available through SuperGrok and the API, and carry forward the always-on reasoning approach Grok 4 introduced. ChatGPT and Claude remain the two most common alternatives for reasoning and coding work respectively, while Gemini pairs well with Google Search and Workspace for research-heavy tasks.
For an open-weight option instead of a hosted subscription, our Grok vs DeepSeek comparison covers DeepSeek’s R1 and V3 models directly against Grok’s lineup. If reasoning-model retirement timelines matter to your decision, OpenAI o1 vs o3 walks through a similar generational handoff on the OpenAI side.
FAQ
Is Grok 4 better than Grok 3?
Yes, on every published benchmark. Grok 4 scored 88% on GPQA Diamond against Grok 3’s 75.4%, and 94% on AIME 2024 against Grok 3’s 52.2% in standard mode. Grok 4 also doubled Grok 3’s shipped context window, from 128K to 256K tokens.
Can I still use Grok 3 in 2026?
Not through the original API. xAI retired the grok-3 endpoint on May 15, 2026, and redirected existing integrations to grok-4.3. The consumer Grok app’s free tier now runs on a newer baseline model instead of Grok 3.
Is Grok 4 still available, or has it also been retired?
The original grok-4-0709 endpoint was retired, but the Grok 4 generation continues under newer names. xAI’s current lineup includes Grok 4.3, Grok 4.5, and Grok 4.6, all built on the same always-on reasoning approach Grok 4 introduced in July 2025.
How much does Grok 4 cost compared to Grok 3?
At launch, third-party trackers listed both models at roughly the same API rate, about $3 per million input tokens and $15 per million output tokens. Grok 4’s always-on reasoning meant it typically used more output tokens per answer than Grok 3’s standard mode, making real-world cost higher even at a matching headline rate.
Is SuperGrok worth $30 a month for Grok 4-tier access?
It depends on how often your work needs deep reasoning. SuperGrok’s $30 tier is worth it if you regularly hit tasks that benefit from a 128K-plus context window and step-by-step reasoning; casual chat users may find the free tier’s Grok 4.3 baseline sufficient.
What is the difference between Grok 4 and Grok 4 Heavy?
Grok 4 Heavy runs multiple reasoning instances in parallel and compares their answers before responding, while standard Grok 4 runs a single reasoning pass. Heavy scored higher on hard benchmarks like Humanity’s Last Exam (50.7% versus Grok 4’s non-tool score in the mid-20s) but requires the $300-a-month SuperGrok Heavy tier.
Final Verdict
Grok 4 was the clear technical upgrade over Grok 3, with large, consistently reproducible benchmark gains in math, science, and coding, plus a genuinely larger context window and real multimodal tool use. Grok 3’s advantage was never raw capability; it was speed, simplicity, and a lower cost per simple answer for people who did not need deep reasoning on every message.
Both points are now mostly historical, since xAI retired both original API endpoints on May 15, 2026, and current usage runs through Grok 4.3, 4.5, or 4.6 instead. If you came here deciding between Grok 3 and Grok 4 specifically, the honest answer is to start with a current SuperGrok subscription and treat this comparison as the reasoning-on-demand-versus-always-on framework that still explains how xAI’s models behave today.
Arslan Abid
AI tools reviewer · AIComparison.ai
Arslan has tracked xAI’s Grok releases since Grok 3’s February 2025 launch, cross-checking benchmark claims against official documentation for every update. Last reviewed: September 2026.