SpaceXAI (formerly xAI) · AI model

Grok 4 Heavy

SupersededReleased: 9 July 2025Access: SuperGrok Heavy only ($300/month)Humanity's Last Exam: 44.4% with tools (xAI)API: Not availableStatus: Superseded by newer Heavy modes

Grok 4 Heavy was xAI's premium, multi-agent version of Grok 4, sold only inside the $300-a-month SuperGrok Heavy plan. It launched on 9 July 2025. Instead of one line of reasoning, it ran several agents in parallel that each explored the problem and then compared answers. xAI called this parallel test-time compute.

On xAI's figures it scored 44.4% on Humanity's Last Exam with tools and 50.7% on the text-only part, the first model to pass 50% there. It was never offered in the API. The Heavy name lives on: SuperGrok Heavy still costs $300 a month and now runs newer multi-agent models, including a 16-agent version of Grok 4.20 in 2026.

5.5/10
Our verdictExpert score 5.5/10

Grok 4 Heavy showed that running many agents at once can lift scores on very hard tests, but it was slow, costly and hard to justify for most people.

What it did well, on xAI's figures:

  • Hardest exams. 44.4% on Humanity's Last Exam with tools, against 38.6% for standard Grok 4, and 50.7% on the text-only part.
  • Olympiad maths. 61.9% on the USAMO 2025 proof competition.
  • Checks itself. Multiple agents compared hypotheses before answering.

What held it back:

  • Price. $300 a month was the most expensive consumer AI plan from a major lab at launch.
  • Speed. DataCamp and others found it much slower than standard Grok 4.
  • No API. Developers could not build on it.
  • Same safety concerns as Grok 4, which shipped with no safety report.

Who should consider Heavy today: researchers and analysts who hit limits on SuperGrok and need the largest agent team. Who should not: almost everyone else. Grok 4.7 through the API, or a $20 to $30 plan from a rival, covers most needs.

Score breakdown

Reasoning8.0
Research depth7.5
Speed3.5
Value3.5
Current relevance3.0

Best for

  • Very hard research, maths and analysis questions
  • Power users already paying for SuperGrok Heavy
  • Getting early access to new Grok models (Grok 4.3 beta went to Heavy first)

What we like

  • Top-tier reasoning scores in mid-2025: 44.4% on Humanity's Last Exam with tools
  • First model to pass 50% on HLE's text-only subset (xAI)
  • Agents cross-check each other, which can catch errors
  • Heavy plan now includes X Premium+ and the highest Grok limits

Watch out for

  • $300 a month, ten times the standard SuperGrok plan
  • Slow: complex answers could take minutes
  • Never available in the API
  • Inherited Grok 4's safety gaps and missing launch safety report

Specs at a glance

DeveloperxAI (now SpaceXAI)
Release date9 July 2025
Base modelGrok 4
MethodSeveral agents work on the problem in parallel and compare results (parallel test-time compute)
SpeedMuch slower than standard Grok 4; hard questions could take several minutes
Context windowNot separately published
AccessGrok app and grok.com with SuperGrok Heavy; no API
Plan price$300/month (SuperGrok Heavy)
Current Heavy planSuperGrok Heavy, $300/month: largest agent team, highest limits, X Premium+ included
SuccessorsGrok 4.20 Heavy (16 agents, 2026) and later Heavy modes

Benchmarks

Benchmarks are standard tests. Vendor-run results are marked as such; independent results are preferred where they exist.

BenchmarkScoreSourceNote
Humanity's Last Exam (with tools)44.4%xAI via Scientific AmericanStandard Grok 4: 38.6%
Humanity's Last Exam (text-only subset)50.7%xAIxAI said it was the first to pass 50%
USAMO 202561.9%xAIOlympiad maths proofs

Pricing

Plan / tierPriceNotes
SuperGrok Heavy (monthly)$300/monthIncluded Grok 4 Heavy at launch; now the top Grok plan
SuperGrok Heavy (annual)$3,000/yearPer third-party plan trackers, September 2026
APINot offered

How Heavy mode works

A normal reasoning model follows one chain of thought. Heavy mode starts several agents on the same question at the same time. Each explores its own approach, then the results are compared and combined. xAI said this lets Grok consider multiple hypotheses at once.

The trade-off is simple: more agents means more compute, more time and more cost for each answer. That is why xAI put it behind a $300 plan instead of offering it widely.

Grok 4 Grok 4 Heavy
Agents 1 Several in parallel
Humanity's Last Exam (tools) 38.6% 44.4%
Speed Normal Much slower
Access SuperGrok, X Premium+, API SuperGrok Heavy only

Heavy mode after Grok 4

xAI kept the Heavy tier as newer models arrived:

  • Grok 4.20 Heavy (March 2026): a 16-agent version of the multi-agent Grok 4.20 for SuperGrok Heavy subscribers.
  • Grok 4.3 beta (17 April 2026): reached SuperGrok Heavy users about two weeks before anyone else.
  • Today: grok.com lists SuperGrok Heavy as the plan with the largest team of collaborating agents, the highest usage at the fastest speed, dedicated support, and X Premium+ at no extra cost.

So Grok 4 Heavy itself is superseded, but the Heavy idea is now a standing feature of the top plan. For plan details see our Grok app review.

Safety

Grok 4 Heavy was built on Grok 4, which launched without a safety report and drew public criticism from researchers at OpenAI and Anthropic. xAI's later Grok 4 model card, published 20 August 2025, did not give separate safety results for Heavy that we could find. See our SpaceXAI page for the company's wider record.

Expert tips
  1. Before paying $300, try the same hard questions on SuperGrok ($30) or SuperGrok Plus ($100). Upgrade only if you keep hitting limits or the answers clearly improve.
  2. Save Heavy mode for questions with a checkable answer, like maths or data analysis. For simple chat it adds wait time without better results.
  3. Heavy subscribers often get new Grok models first, as with the Grok 4.3 beta. If early access matters to your work, factor that in.
  4. If you already pay for X Premium+, note that SuperGrok Heavy includes it, so you can cancel the separate X subscription.

Jargon explained

Test-time compute
Extra computing power a model uses while answering, for example by thinking longer or running several attempts.
Agent
An AI that works through steps on its own, such as searching, calculating and checking, instead of giving one quick reply.
USAMO
The USA Mathematical Olympiad, a proof-based competition for top high-school maths students.

Alternatives to consider

Frequently asked questions

How much did Grok 4 Heavy cost?

It came only with SuperGrok Heavy, at $300 a month. That plan still costs $300 a month, or $3,000 a year, as of September 2026.

Is Grok 4 Heavy in the API?

No. xAI never offered Grok 4 Heavy through the API. Developers who want multi-agent answers can use the Grok 4.20 multi-agent API model instead.

What is the difference between Grok 4 and Grok 4 Heavy?

Grok 4 Heavy runs several agents on your question at once and compares their answers. It scored higher on hard tests (44.4% vs 38.6% on Humanity's Last Exam with tools) but was much slower and far more expensive.

Is SuperGrok Heavy worth $300 a month?

For most people, no. It makes sense only if you regularly run very hard research or analysis and keep hitting limits on cheaper plans.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Grok 4 (xAI)
  2. Elon Musk's New Grok 4 Takes on Humanity's Last Exam as the AI Race Heats Up (Scientific American)
  3. Grok 4: tests, features, benchmarks, access (DataCamp)
  4. xAI launches Grok 4 along SuperGrok Heavy, a $300 premium subscription (Data Phoenix)
  5. Grok plans (SpaceXAI)
  6. Grok 4.20 multi-agent system: how the 4 agents work (Verdent)
  7. xAI rolls out Grok 4.3 beta for SuperGrok Heavy subscribers (PiunikaWeb)
  8. Grok pricing 2026: plans and API costs (AI Toolbox)

More from SpaceXAI (formerly xAI)