Elon Musk's AI company xAI has released Grok 4.7, its new flagship model for coding, agentic tasks and knowledge work. In the announcement covered by MarkTechPost the company appears under the name SpaceXAI. The headline is the price: $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6.
The real story is the gap between two tables: the company's own numbers show a big leap, while independent measurement puts the model in the middle of the pack.
What changed
xAI lists four changes over Grok 4.6:
- A new, larger base model; Grok 4.6's base was not reused.
- A longer reinforcement learning run weighted toward hard tasks that take hours to complete.
- More careful self-verification and better long-context handling.
- Native support for the Grok Bot harness for conversational and knowledge work.
Two tables, two verdicts
On the company's chart, Grok 4.7 beats Grok 4.6 in every row. The biggest jump is on Terminal-Bench 4.0, from 20.3% to 38%. On Harvey's legal agent benchmark it scores 19.6% against 6.7% for Claude Fable 5.1 Max. Even so, the model does not lead everywhere on its own chart: Fable 5.1 Max tops four of the seven benchmarks.
The independent Artificial Analysis index is cooler. As The Decoder reports, on the index that combines ten benchmarks Grok 4.7 scores 46, while Claude Fable 5.1 and GPT-6 lead with 53 each.
| Terminal-Bench 4.0 | Score |
|---|---|
| Grok 4.7, xAI's own measurement | 38% |
| Grok 4.7, independent | 26% |
| GPT-6 Astra, independent | 60% |
| Claude Fable 5.1, independent | 55% |
| DeepSeek V4.1 Flash, independent | 27% |
Part of the gap may come down to settings: xAI reports its results at the highest effort level (xHigh), and all of its scores are vendor-reported. Still, it is notable that a far cheaper Chinese model, DeepSeek V4.1 Flash, edges past Grok 4.7 in the independent run.
The price argument
xAI's case rests on cost. By its numbers, Fable 5.1 Max is 5 times more expensive on input and about 8.3 times on output. The Decoder notes that these rates sit closer to Chinese models than to Western frontier models, probably for good reason. On a cost-per-task chart, xAI places the model on the price-performance frontier.
The model is available through the xAI API, on all Cursor plans, in Grok Build, and via OpenRouter, Vercel and Cloudflare. Grok 4.7 Fast, with twice the output speed at twice the price, runs only in Cursor and Grok Build. A US regional endpoint that keeps inference in the United States costs 10% more.
Safety
Grok 4.7 ships with an entirely new safeguard stack. The company says it topped LatchBio's biosafety benchmark with 62.4%. On its own HackerBench test, it says the model let 3.3% of risky dual-use cyber prompts through.
For developers the summary is simple: Grok 4.7 is cheap and clearly better than its predecessor, but it has not closed the gap to the top models. Before choosing, look at cost per task on your own workload rather than the vendor's chart.