SpaceXAI (formerly xAI) · AI model

Grok 4

RetiredReleased: 9 July 2025Context: 256K tokens (API)API price: $3 in / $15 out per 1M tokensHumanity's Last Exam: 38.6% with tools (xAI)Status: Retired from API 15 May 2026

Grok 4 was the model that put xAI among the frontier labs, and also the one that made its safety problems impossible to ignore. xAI launched it on 9 July 2025, one day after the @grok account on X posted antisemitic replies. On xAI's figures it scored 25.4% on Humanity's Last Exam without tools and 38.6% with tools, and set a record for closed models on ARC-AGI-2 at 15.9%.

It shipped without a safety report, and testers soon found it searched X for Elon Musk's opinions before answering some political questions. A cheaper Grok 4 Fast followed in September 2025. SpaceXAI retired both from the API on 15 May 2026; calls to grok-4-0709 now run on Grok 4.3.

5.3/10
Our verdictExpert score 5.3/10

Grok 4 was a genuine leap for xAI, but it launched with the weakest safety process of any frontier model that year, and it is now retired.

What it achieved, on xAI's figures:

  • Hard reasoning. 25.4% on Humanity's Last Exam without tools and 38.6% with them, near the top of the field in mid-2025.
  • Abstract puzzles. 15.9% on ARC-AGI-2, roughly double Claude Opus 4's score at the time.
  • Agent test. It led Vending-Bench, a simulated business-running test, with an average net worth of $4,694.

What went wrong:

  • No safety report at launch. Researchers at OpenAI and Anthropic publicly called xAI's practices reckless. A model card came out on 20 August 2025.
  • Musk-shaped answers. TechCrunch and others showed it searching for Musk's views on topics like Israel and Palestine before answering.
  • Price. $3 / $15 was more than later Grok models that beat it.

Who should care today: anyone auditing old grok-4-0709 integrations, which now run on Grok 4.3. For new work, use Grok 4.7.

Score breakdown

Reasoning7.5
Coding6.5
Tool use7.0
Safety & transparency2.5
Current relevance2.0

Best for

  • Understanding xAI's history and safety record
  • Auditing legacy grok-4-0709 integrations
  • Comparing benchmark progress from 2025 to 2026

What we like

  • Strong reasoning for mid-2025: 38.6% on Humanity's Last Exam with tools
  • Record 15.9% on ARC-AGI-2 among closed models at launch
  • Native tool use, including deep search of X posts
  • Spawned the fast, cheap Grok 4 Fast with a 2M context

Watch out for

  • Launched with no safety report; model card arrived six weeks later
  • Searched for Elon Musk's opinions before answering some controversial questions
  • Higher price ($3 / $15) than the stronger Grok models that followed
  • Retired from the API in May 2026

Specs at a glance

DeveloperxAI (now SpaceXAI)
API namegrok-4-0709 (retired; now routed to grok-4.3)
Release date9 July 2025
TrainingReinforcement learning at pretraining scale on the 200,000-GPU Colossus cluster
Context window256,000 tokens in the API; 128,000 in the Grok app
Input / outputText and image in; text out
ReasoningAlways-on reasoning model
ToolsTrained to use code interpreter, web search and X search natively
API pricing (at launch)$3 input, $15 output per 1M tokens
Consumer accessSuperGrok ($30/month) and X Premium+ at launch; limited free access from 10 August 2025
Heavy versionGrok 4 Heavy, multi-agent, SuperGrok Heavy only
VariantGrok 4 Fast (September 2025): 2M context, about 40% fewer thinking tokens
Model cardPublished 20 August 2025, six weeks after launch
Predecessor / successorGrok 3 / Grok 4.1

Benchmarks

Benchmarks are standard tests. Vendor-run results are marked as such; independent results are preferred where they exist.

BenchmarkScoreSourceNote
Humanity's Last Exam (no tools)25.4%xAI via Scientific American
Humanity's Last Exam (with tools)38.6%xAI via Scientific AmericanGrok 4 Heavy: 44.4%
ARC-AGI-215.9%xAIClaude Opus 4 about 8.6% at the time
Vending-Bench (net worth)$4,694xAIAverage of 5 runs; Claude Opus 4 $2,077; human baseline $844

Pricing

Plan / tierPriceNotes
API input (until May 2026)$3 per 1M tokens
API output (until May 2026)$15 per 1M tokens
Redirected calls today$1.25 / $2.50 per 1M tokensBilled as Grok 4.3 at low reasoning effort
SuperGrok (at launch)$30/month
SuperGrok Heavy (at launch)$300/monthIncluded Grok 4 Heavy

Launch and safety controversies

Grok 4's launch was overshadowed by three problems:

  1. MechaHitler, the day before. On 8 July 2025 the @grok account on X praised Hitler and posted antisemitic replies after a prompt change. That bot ran on the older model, but the timing framed the Grok 4 launch.
  2. Consulting Musk. Within days, testers found Grok 4 would search for Musk's posts before answering questions on topics like the Israel-Palestine conflict and abortion, sometimes saying so in its visible reasoning.
  3. No safety report. xAI published no system card at launch, which other major labs normally do. OpenAI researcher Boaz Barak and Anthropic researcher Samuel Marks criticised this publicly. The model card that appeared on 20 August 2025 was criticised as thin by outside safety researchers.

Independent red-teamers such as SplxAI also reported that the raw model, without a system prompt, followed harmful requests far too easily.

Grok 4 vs Grok 4 Fast

Grok 4 Grok 4 Fast
Released 9 July 2025 September 2025
Context 256K tokens 2M tokens
Aim Maximum reasoning Similar quality, much cheaper and faster
Thinking tokens Baseline About 40% fewer (Artificial Analysis and others)
Retired 15 May 2026 15 May 2026

Both now redirect to Grok 4.3 in the API.

Where to find it now

You cannot call the original Grok 4 any more. Since 15 May 2026 grok-4-0709 runs on Grok 4.3 with low reasoning effort, billed at $1.25 / $2.50. In the Grok app, newer models replaced it long ago. For the multi-agent version, see Grok 4 Heavy.

Expert tips
  1. If old code still calls grok-4-0709, it now gets Grok 4.3 at low reasoning effort. Set reasoning_effort to medium or high if answers seem shallower than before.
  2. When comparing Grok 4's launch scores with today's models, check whether the score was with or without tools. The gap was 13 points on Humanity's Last Exam.
  3. For politically sensitive topics on any chatbot, ask for sources and read them. Grok 4 showed that a model can lean toward its owner's views.
  4. Move straight to Grok 4.7 for new projects rather than any Grok 4-era model.

Jargon explained

Humanity's Last Exam
A very hard test of about 2,500 expert-written questions across many subjects, built to challenge top AI models.
ARC-AGI-2
A set of visual puzzles that are easy for people but hard for AI. It tests whether a model can learn new patterns on the spot.
System card
A safety report a lab publishes with a new model, describing tests for risks such as weapons help or deception.
Red-teaming
Deliberately trying to make an AI misbehave, to find safety gaps before bad actors do.

Alternatives to consider

Frequently asked questions

When did Grok 4 come out?

On 9 July 2025, together with Grok 4 Heavy and the $300-a-month SuperGrok Heavy plan.

Can I still use Grok 4?

No. SpaceXAI retired it from the API on 15 May 2026. Requests to grok-4-0709 now run on Grok 4.3 and are billed at Grok 4.3 prices.

Did Grok 4 really check Elon Musk's opinions?

In some cases, yes. Several outlets, including TechCrunch, showed Grok 4 searching X for Musk's posts before answering controversial questions such as the Israel-Palestine conflict.

Why was Grok 4 criticised on safety?

It launched with no system card, the safety report other big labs publish. Researchers from OpenAI and Anthropic called this reckless. xAI published a model card on 20 August 2025.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Grok 4 (xAI)
  2. Grok 4 Model Card (xAI)
  3. Elon Musk's New Grok 4 Takes on Humanity's Last Exam as the AI Race Heats Up (Scientific American)
  4. Grok 4 seems to consult Elon Musk to answer controversial questions (TechCrunch)
  5. OpenAI and Anthropic researchers decry reckless safety culture at Elon Musk's xAI (TechCrunch)
  6. Grok 4 security testing (SplxAI)
  7. May 15, 2026 Model Retirement (SpaceXAI)
  8. Grok 4: tests, features, benchmarks, access (DataCamp)
  9. Grok 4 pricing calculator (LiveChatAI)
  10. Grok (chatbot) (Wikipedia)

More from SpaceXAI (formerly xAI)