SpaceXAI (formerly xAI) · AI model

Grok 4.20 (Grok 4.2)

SupersededReleased: 17 Feb 2026 (app beta); 10 Mar 2026 (API)Context: 2M tokens at launch; 1M listed for the multi-agent model nowAPI price now: $1.25 in / $2.50 out per 1M tokens (multi-agent)Agents: 4 in the app; 16 in Heavy modeStatus: Superseded by Grok 4.3 (Apr 2026)

Grok 4.20, often written Grok 4.2, was the Grok that answered with a team of agents instead of one. It launched as a public beta in the Grok app on 17 February 2026 and reached the xAI API on 10 March 2026. In the app, four named agents (Grok, Harper, Benjamin and Lucas) worked on hard questions in parallel, checked each other, and agreed a final answer. SuperGrok Heavy subscribers got a larger 16-agent version.

SpaceXAI also said the beta would be updated weekly based on real use. Artificial Analysis scored it 49 on its Intelligence Index (on the older, pre-v4.3 scale). Just ten weeks later, Grok 4.3 beat it by 4 points at a much lower price. The multi-agent API model is still available.

6.5/10
Our verdictExpert score 6.5/10

Grok 4.20 was an interesting experiment in making AI agents argue before they answer, but Grok 4.3 and later models have overtaken it.

What it brought:

  • Built-in debate. Instead of one chain of thought, several agents tackled a question and one agent's job was to disagree. The goal was fewer confident mistakes.
  • Huge context at launch. Third-party listings recorded a 2-million-token window.
  • Research focus. SpaceXAI still describes the multi-agent API model as a tool for deep research tasks.

What held it back:

  • Cost and speed. Many agents mean more compute per answer, and at launch it cost $2 / $6.
  • Quickly beaten. Grok 4.3 scored 4 points higher on the Artificial Analysis index and cut the price by 37.5% on input and 58% on output.
  • Moving target. Weekly updates made it hard to reproduce results.

Who should use it: researchers who want the multi-agent API model for long, open-ended research questions. Who should not: anyone who needs a stable, cheap general model; use Grok 4.3 or Grok 4.7.

Score breakdown

Reasoning7.2
Coding6.3
Research & multi-agent7.6
Value7.0
Current relevance4.5

Best for

  • Deep research questions where you want several viewpoints
  • Experimenting with multi-agent answers via the API
  • Batch research jobs at low cost

What we like

  • Agents check each other's work before answering
  • Heavy mode runs 16 agents for the hardest questions
  • Multi-agent API model now costs just $1.25 / $2.50 with batch support
  • Very large context window at launch

Watch out for

  • Beaten by Grok 4.3 within ten weeks on both score and price
  • Weekly behaviour changes during the beta made results hard to repeat
  • No log probabilities in the API, which some evaluation tools need
  • Launched in the middle of Grok's deepfake scandal and regulatory probes

Specs at a glance

DeveloperxAI (now SpaceXAI)
Official nameGrok 4.20; widely called Grok 4.2
API namesgrok-4.20 and grok-4.20-multi-agent-0309 (aliases include grok-4.20-multi-agent and grok-4.20-multi-agent-beta-0309)
Release datesApp beta: 17 February 2026; API: 10 March 2026
How it worksSeveral agents reason in parallel, critique each other and merge a final answer
App agentsGrok (coordinator), Harper, Benjamin and Lucas (the designated contrarian)
Heavy mode16 agents, for SuperGrok Heavy subscribers
Input / outputText and image in; text out
Context window2M tokens at launch (third-party listings); 1M on the current multi-agent model page
API pricing at launch$2 input, $6 output per 1M tokens
API pricing now (multi-agent)$1.25 input, $0.20 cached, $2.50 output per 1M tokens; batch supported
Log probabilitiesNot supported from Grok 4.20 onward
Predecessor / successorGrok 4.1 / Grok 4.3

Benchmarks

Benchmarks are standard tests. Vendor-run results are marked as such; independent results are preferred where they exist.

BenchmarkScoreSourceNote
Artificial Analysis Intelligence Index (Apr 2026)49Artificial AnalysisOn the older (pre-v4.3) scale; not comparable with current v4.3 scores. Grok 4.3 scored 53, 4 points higher
GDPval-AA (Elo)About 1,179Artificial AnalysisDerived: Grok 4.3's 1,500 was 321 points higher
tau2-Bench TelecomAbout 93%Artificial AnalysisDerived: Grok 4.3's 98% was 5 points higher

Pricing

Plan / tierPriceNotes
API input (multi-agent, now)$1.25 per 1M tokensHigher above 200K prompt tokens
API cached input$0.20 per 1M tokens
API output (multi-agent, now)$2.50 per 1M tokens
API at launch$2 / $6 per 1M tokensInput / output
Heavy mode in the appSuperGrok Heavy, $300/month

How the four agents worked

In the Grok app, a hard question went to four agents at once:

Agent Role described at launch
Grok Coordinates the team and writes the final answer
Harper Research and fact gathering
Benjamin Logic, maths and code checks
Lucas Argues against the others to catch mistakes

The roles come from launch coverage, not a technical paper from SpaceXAI. The idea is similar to asking several experts and only accepting what survives their debate. It costs more compute per answer, which is why Heavy mode, with 16 agents, sat on the $300 plan.

Rapid-learning beta

SpaceXAI said Grok 4.20 would be updated weekly using feedback from real conversations. That let it fix problems quickly, but it also meant the model you tested one week might behave differently the next. For businesses this is a real downside. If you use the API, pin a dated name such as grok-4.20-multi-agent-0309 rather than a -latest alias.

Context and safety

Grok 4.20 arrived while Grok was under investigation in the UK and EU over sexualised deepfakes made with its image tools on X (see our SpaceXAI page). Those problems came from Grok's image features, not from the 4.20 text model itself. We found no separate model card or safety report for Grok 4.20.

Expert tips
  1. Use the multi-agent model for research questions, not simple chat. It spends more compute per answer.
  2. Pin grok-4.20-multi-agent-0309 in production so weekly beta updates cannot change behaviour under you.
  3. If your evaluation setup relies on logprobs, it will not work with Grok 4.20 or newer; the field is silently ignored.
  4. For general tasks, switch to Grok 4.3 at the same current price; it scored higher on Artificial Analysis's index.

Jargon explained

Multi-agent
Several copies or roles of an AI working on the same task at once and combining their results.
Log probabilities (logprobs)
Numbers showing how confident a model was in each word it chose. Some testing tools use them.
Beta
An early public version that may change often and may have bugs.

Alternatives to consider

Frequently asked questions

Is it Grok 4.2 or Grok 4.20?

SpaceXAI's official name is Grok 4.20, and the API model names use 4.20. Many people and articles call it Grok 4.2.

When was Grok 4.20 released?

As a public beta in the Grok app on 17 February 2026, and in the xAI API on 10 March 2026.

What are Harper, Benjamin and Lucas?

They are the names of three of the four agents that worked together in Grok 4.20's app version, with Grok as the coordinator. Lucas's job was to disagree with the others.

Can I still use Grok 4.20?

Yes. The multi-agent model is still listed in the xAI API at $1.25 / $2.50 per million tokens as of September 2026.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Grok 4.20 Multi Agent Beta 0309 model page (SpaceXAI)
  2. Release Notes (SpaceXAI)
  3. Grok Models & Pricing (SpaceXAI)
  4. Grok-4.20 Multi-Agent Beta benchmarks, pricing and context window (LLM Stats)
  5. xAI launches Grok 4.20 and it has 4 AI agents collaborating (NextBigFuture)
  6. Grok 4.20 multi-agent system: how the 4 agents work (Verdent)
  7. xAI launches Grok 4.3 with improved agentic performance and lower pricing (Artificial Analysis)

More from SpaceXAI (formerly xAI)