SpaceXAI (formerly xAI) · AI model

Grok 4.7

CurrentReleased: 21 September 2026Context: 500K tokensAPI price: $2 in / $6 out per 1M tokensKnowledge cutoff: May 2026AA Intelligence Index: 46 (Sep 2026)

Grok 4.7 is SpaceXAI's best model and one of the cheapest near-frontier models you can buy, but it is not the smartest. It came out on 21 September 2026 in Cursor, Grok Build, the Grok app and the xAI API. It costs $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6.

SpaceXAI says it uses a new, larger base model and a longer training run focused on tasks that take hours. Decrypt reported it has about 2.1 trillion parameters, up from about 1.5 trillion. On SpaceXAI's own table it beats GPT-5.6 Sol on CursorBench 4.0 (46.3% vs 41.7%) but trails Claude Fable 5.1 (51.8%). Independent group Artificial Analysis scores it 46 on its Intelligence Index, behind GPT-6 Astra and Claude Fable 5.1 (both 53).

7.8/10
Our verdictExpert score 7.8/10

Grok 4.7 is the best value model for long coding and office tasks if you are comfortable with SpaceXAI as a vendor.

The case for it:

  • Price. At $2 / $6 it costs half of GPT-5.6 Sol's input price and under a third of its output price, and a fifth of Claude Fable 5.1 on input.
  • Real gains over Grok 4.6. SpaceXAI reports CursorBench 4.0 up from 40.4% to 46.3%, Terminal-Bench 4.0 up from 20.3% to 37.6% (independent testing by Artificial Analysis, reported by The Decoder, found 26%), and HealthBench Professional up from 48.5% to 56.7%.
  • Speed. Artificial Analysis measured about 188 tokens per second.

The case against:

  • Still a rung below the top. It scores 46 on the Artificial Analysis Intelligence Index, against 53 for GPT-6 Astra and Claude Fable 5.1. Fable 5.1 also beats it on Terminal-Bench 4.0 (57.9%).
  • Wordy. Artificial Analysis counted about 81,000 output tokens per index task, which eats into the low price.
  • Trust. SpaceXAI claims a new safeguard stack, but Grok's past failures mean you should test it yourself.

Choose it for high-volume coding agents, Cursor users, and cost-sensitive document work. Skip it if you need the single strongest model, or a vendor with a clean safety record.

Score breakdown

Reasoning8.2
Coding8.4
Agentic & computer use8.2
Value9.3
Safety & trust5.0

Best for

  • Coding agents in Cursor or Grok Build
  • High-volume agent pipelines where cost per task matters
  • Drafting documents, spreadsheets and slides
  • Teams already using Grok 4.6 who want a free upgrade

What we like

  • Low price for its class: $2 in / $6 out per 1M tokens
  • Big jump on long terminal tasks: 37.6% on Terminal-Bench 4.0 vs 20.3% for Grok 4.6 (SpaceXAI's figure; independent testing found 26%)
  • Strong on legal and electrical-engineering tests in SpaceXAI's table (Harvey LAB 19.6%, EEBench 64.0%)
  • Available on day one in Cursor, GitHub Copilot and the API
  • 500K-token context with no fixed output cap

Watch out for

  • Trails GPT-6 Astra and Claude Fable 5.1 on independent scoring (46 vs 53)
  • Uses many output tokens per task, so real bills can run higher than the list price suggests
  • Smaller context than Grok 4.3's 1M tokens; prices double above 200K
  • Most benchmark figures come from SpaceXAI itself

Specs at a glance

DeveloperSpaceXAI (formerly xAI)
API namegrok-4.7
Release date21 September 2026
SizeAbout 2.1 trillion parameters (reported by Decrypt; not in the official post)
Context window500,000 tokens
Output limitNo text output limit, per SpaceXAI release notes
Input / outputText and image in; text out
Knowledge cutoffMay 2026
Reasoning effortlow, medium, high (default), xhigh
API pricing (under 200K prompt)$2 input, $0.50 cached input, $6 output per 1M tokens
API pricing (200K+ prompt)$4 input, $1 cached input, $12 output per 1M tokens
Fast variantAbout 2x output speed at 2x the price
Batch APINot supported
Where to use itGrok app, Cursor, Grok Build, xAI API, GitHub Copilot, model routers
PredecessorGrok 4.6 (12 August 2026)

Benchmarks

Benchmarks are standard tests. Vendor-run results are marked as such; independent results are preferred where they exist.

BenchmarkScoreSourceNote
Artificial Analysis Intelligence Index v4.346Artificial Analysis (via OfficeChai)GPT-6 Astra 53, Claude Fable 5.1 53, GPT-5.6 Sol 47, Grok 4.6 44
CursorBench 4.046.3%SpaceXAIGrok 4.6 40.4%; GPT-5.6 Sol 41.7%; Fable 5.1 51.8%
DeepSWE v1.171.0% (high effort)SpaceXAIGPT-5.6 Sol 72.7%; Fable 5.1 70.0%
Terminal-Bench 4.037.6%SpaceXAIVendor figure. Artificial Analysis measured 26% (reported by The Decoder). Grok 4.6 20.3%; Fable 5.1 57.9% in SpaceXAI's table
AA-Briefcase v1.1 (Elo)1,657SpaceXAI / Artificial AnalysisFable 5.1 1,678; Grok 4.6 1,546
HealthBench Professional56.7%SpaceXAIFable 5.1 62.1%; GPT-5.6 Sol 60.5%
Harvey Legal Agent Benchmark19.6%SpaceXAIGrok 4.6 15.8%; Fable 5.1 6.7%
LatchBio biosafety benchmark62.4%SpaceXAIVendor claim: top score

Pricing

Plan / tierPriceNotes
API input$2 per 1M tokens$4 when the prompt is 200K tokens or more
API cached input$0.50 per 1M tokens$1 above 200K
API output$6 per 1M tokens$12 above 200K
Fast variantAbout 2x standard ratesRoughly twice the output speed
Grok appFree tier; SuperGrok from $10/monthHigher plans get more Grok 4.7 usage; see our Grok app page

What changed from Grok 4.6

SpaceXAI lists three main changes:

  1. A new, larger base model. Decrypt reported 2.1 trillion parameters, about 40% more than Grok 4.6.
  2. Longer reinforcement learning (training by trial and reward) on harder tasks, weighted toward problems that take many hours.
  3. Native Grok Bot support. It was trained to work inside SpaceXAI's always-on agent harness.

Decrypt also reported that training included SpaceX data such as Starlink telemetry and engineering failure logs.

Test (SpaceXAI figures) Grok 4.7 Grok 4.6
CursorBench 4.0 46.3% 40.4%
DeepSWE v1.1 71.0% 65.2%
EEBench 64.0% 53.0%
Terminal-Bench 4.0 37.6% (independent: 26%) 20.3%
AA-Briefcase v1.1 1,657 1,546
HealthBench Professional 56.7% 48.5%

How it compares with GPT and Claude

On SpaceXAI's own comparison table:

Grok 4.7 GPT-5.6 Sol Claude Fable 5.1
Price (in / out per 1M) $2 / $6 $4 / $20 $10 / $50
CursorBench 4.0 46.3% 41.7% 51.8%
DeepSWE v1.1 71.0% 72.7% 70.0%
Terminal-Bench 4.0 37.6% 37.3% 57.9%
HealthBench Professional 56.7% 60.5% 62.1%

The pattern is clear. Grok 4.7 matches or beats GPT-5.6 Sol on coding for less money, but Claude Fable 5.1 is stronger on long terminal work and medicine. Artificial Analysis puts GPT-6 Astra and Claude Fable 5.1 seven points ahead on its index.

Safety

SpaceXAI says Grok 4.7 has an entirely new safeguard stack and is its strongest model yet at refusing jailbreaks (tricks to get around safety rules). It claims the model lets only 3.3% of risky dual-use cyber prompts through on its own HackerBench v0.3 test, and it has given some security partners invite-only access to its red-team abilities.

These are vendor claims. At publication we had not found an independent safety evaluation or a separate model card. Given Grok's history, covered on our SpaceXAI page, treat them with care until outside testers confirm them.

Expert tips
  1. Start at high reasoning effort (the default) and only use xhigh for tasks that fail at high. Artificial Analysis found the model already uses about 81,000 output tokens per hard task.
  2. Keep prompts under 200K tokens. Crossing that line doubles the price of every token in the request.
  3. Use cached input for repeated system prompts and codebases. It drops input cost from $2 to $0.50 per million tokens.
  4. On the Responses API, Grok 4.7 always returns encrypted reasoning content. Budget for the extra payload size if you log full responses.
  5. Pin grok-4.7 rather than a -latest alias in production so a future update does not change behaviour without warning.

Jargon explained

Reasoning effort
A setting that controls how long the model thinks before answering. Higher effort is slower and costs more but can be more accurate.
CursorBench
A coding test made by Cursor that measures how well a model finishes real, multi-step programming tasks in the Cursor editor.
Terminal-Bench
A test of how well an AI can complete jobs by typing commands in a computer terminal.
Cached input
Text the provider has seen recently in your requests, such as a long system prompt. It is billed at a lower rate the second time.

Alternatives to consider

Frequently asked questions

When was Grok 4.7 released?

On 21 September 2026, in the Grok app, Cursor, Grok Build and the xAI API at the same time.

How much does Grok 4.7 cost?

$2 per million input tokens and $6 per million output tokens, with cached input at $0.50. Prompts of 200,000 tokens or more cost double. A fast variant costs about twice as much.

Is Grok 4.7 better than GPT-6 or Claude?

Not overall. Artificial Analysis scores it 46, against 53 for GPT-6 Astra and Claude Fable 5.1. It does beat GPT-5.6 Sol on SpaceXAI's CursorBench 4.0 test and is much cheaper than both rivals.

Can I use Grok 4.7 for free?

SpaceXAI offered free use in Grok Build at launch, and the Grok app has a limited free tier. Heavier use needs a SuperGrok plan or API credits.

What is Grok 4.7's context window?

500,000 tokens, which is roughly 375,000 English words. That is half of Grok 4.3's 1 million tokens.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Introducing Grok 4.7 (SpaceXAI)
  2. Grok 4.7 model page (SpaceXAI)
  3. Release Notes (SpaceXAI)
  4. Grok Models & Pricing (SpaceXAI)
  5. xAI Launches Grok 4.7. It's Bigger, But Late to the AI Frontier Party (Decrypt via Yahoo Tech)
  6. Grok 4.7's score jumps 2 points on Artificial Analysis Intelligence Index (OfficeChai)
  7. Grok 4.7 model analysis (Artificial Analysis)
  8. Grok 4.7 is now available in GitHub Copilot (GitHub)
  9. xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6 (The Decoder)

More from SpaceXAI (formerly xAI)