Grok 4.7 arrives at the same price but uses twice the tokens
SpaceXAI's new model, released on Monday, beats its predecessor on every coding benchmark the company lists and keeps the $2 and $6 per-million pricing. Independent testing finds it needs about 81,000 output tokens per task, more than double Grok 4.6.

Key takeaways
- Grok 4.7 costs $2 per million input and $6 per million output tokens, unchanged from Grok 4.6.
- SpaceXAI's table: Terminal-Bench 4.0 up from 20.3% to 38.0%, DeepSWE v1.1 from 65.2% to 71.0%.
- Artificial Analysis measured about 81,000 output tokens per task, against 36,000 for Grok 4.6.
SpaceXAI, the company formerly known as xAI, released Grok 4.7 on Monday, September 21, calling it its most capable model for coding and knowledge work. It is available in Cursor, in the company's own Grok Build coding agent, through the Grok API, in third-party coding harnesses and on the usual model routers and clouds. The price is unchanged from Grok 4.6: $2 per million input tokens and $6 per million output tokens.
The model takes text and images, writes text, has a context window of 500,000 tokens and a knowledge cut-off of May 2026, and can be run at four reasoning-effort levels, from low to "xhigh". A "Fast" variant served on quicker hardware costs double the standard rates and is available only in Cursor and Grok Build, not on the public API.
The numbers SpaceXAI wants you to see
The company's own table shows gains across the board over Grok 4.6, which it says was trained on a smaller base model. Terminal-Bench 4.0 rises from 20.3 per cent to 38.0; DeepSWE v1.1 from 65.2 to 71.0; CursorBench 4.0 from 40.4 to 46.3; EEBench from 53.0 to 64.0; Harvey's legal-agent benchmark from 15.8 to 19.6; HealthBench Professional from 48.5 to 56.7. On GDPval, an Elo-style measure of knowledge work, it moves from 1,605 to 1,695.
Even on that table, Grok 4.7 is not top of the class. SpaceXAI lists Anthropic's Fable 5.1 at 57.9 on Terminal-Bench 4.0 and 51.8 on CursorBench 4.0, well ahead of it, and OpenAI's GPT-5.6 Sol ahead on DeepSWE and HealthBench. The company also says the model is "the strongest model we've tested on refusals and jailbreak resistance", citing 62.4 per cent on LatchBio's biosafety benchmark and a 3.3 per cent pass-through rate for risky dual-use prompts on HackerBench v0.3. Those are company-reported figures; nobody outside has checked them yet.
The cadence is worth a moment. Grok 4.6 shipped on August 12 at the same prices. Forty days later its successor is out, and the announcement's promise is the same as before: "twice as fast, at half the price of comparable models".
The number that decides your bill
Artificial Analysis, which runs its own tests, gives Grok 4.7 a score of 46 on its Intelligence Index, two points above Grok 4.6, and finds the model now sits at the frontier on agentic knowledge work: up 111 Elo on the AA-Briefcase test to 1,657, and 1,695 on GDPval-AA. Paired with Grok Build, it scores 56 on the Coding Agent Index, fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5. Outside agentic knowledge work, AA says, it "broadly matches" Grok 4.6 at high effort. Its hallucination rate on AA's Omniscience test fell to 29 per cent from 34.
Then comes the line that should be in every procurement note. At xhigh effort, Grok 4.7 used roughly 81,000 output tokens per Intelligence Index task, against 36,000 for Grok 4.6 at high effort. Each task took about 7.1 minutes at roughly 188 tokens per second.
Why does that matter when the per-token price has not moved? Because output tokens are what you pay for. Artificial Analysis's cost-per-task figure, reported by VentureBeat, puts Grok 4.7 at about $3.74 per task at xhigh effort and $2.73 at high, against about $1.99 for GPT-5.6 Sol at maximum effort, even though the OpenAI model lists at $4 and $20. As VentureBeat puts it, "a model charging less per token can still be more expensive on a finished workload if it needs substantially more reasoning tokens to get there". Cache hits are discounted to $0.50 per million, which helps on long, repeated contexts but not on reasoning.
What the comparison is really about
Two things can be true. Grok 4.7 is a clear step up from Grok 4.6 on agent and coding tasks, and it is a heavier, slower model to run for the same answer. That suggests the reasoning-effort dial is the setting to test first: the headline benchmarks were run at xhigh, and a team that leaves the default "high" may see neither the full gains nor the full token bill. SpaceXAI has not published token-efficiency figures of its own.
The other thing to watch is the harness. SpaceXAI says it trained the model to "natively understand" its Grok Bot harness, and the Fast variant is confined to its own coding tools and Cursor. That is a company steering usage toward surfaces it controls, which is where model companies increasingly make their margin.
- SpaceXAI
- Grok 4.7
- xAI
- AI models
- coding models
- Artificial Analysis
Sources
- Introducing Grok 4.7 — SpaceXAI, Sep 21, 2026
- Grok 4.7 | SpaceXAI Docs — SpaceXAI
- Introducing Grok 4.6 — SpaceXAI, Aug 12, 2026
- Benchmarking Grok 4.7 — Artificial Analysis, Sep 21, 2026
- Grok 4.7 pairs coding gains with the same affordable pricing — but high token consumption threatens real-world ROI — VentureBeat, Sep 21, 2026
Related stories

Qwen-Image-2.1 weights arrive under a research-only licence
The new 7-billion-parameter image model generates and edits pictures with transparent backgrounds at 2K. You can download it, but you cannot build a business on it without asking.
3 min read
Comments
No comments yet. Start the conversation.