Grok 4.7 launched on September 21, 2026, at $2 per million input tokens and $6 per million output, 80% below Claude Fable 5.1’s input price. xAI says it edges Fable on DeepSWE (71.0% vs 70%), but Artificial Analysis scores it 46 on its Intelligence Index, seven points behind Fable 5.1 and GPT-6 at 53.
Grok 4.7 is finally out. After at least five missed release targets since late July, xAI shipped its new flagship on Sunday and pushed it straight into Cursor, the Grok API and GitHub Copilot on the same day. The pitch is simple: frontier-adjacent coding at a Chinese-model price.
The catch is just as simple. Independent testing puts Grok 4.7 clearly behind the two models that set the enterprise price ceiling. That makes this a pricing story first and a capability story second.
What is Grok 4.7?
Grok 4.7 is xAI’s newest large language model, released September 21, 2026, and tuned for agentic coding and long multi-hour tasks. xAI says it uses a larger base model than Grok 4.6, extended reinforcement learning, better self-verification and a new safeguard stack. It ships in a standard and a 2x-faster variant.
Elon Musk has put the model at 2.1 trillion parameters, up about 40% from Grok 4.6’s 1.5 trillion, according to Decrypt. xAI has not published an official spec sheet, so treat that figure as Musk’s claim, not a disclosed number.
The training mix is the unusual part. Musk said SpaceX’s engineering corpus, excluding ITAR-restricted material, entered “supplemental training of the 2T run,” per CellCog’s release timeline. Decrypt reports that includes Starlink telemetry, manufacturing records and engineering failure logs.
What xAI changed from Grok 4.6
- Bigger base model: roughly 2.1 trillion parameters, by Musk’s count.
- Harder RL: extended reinforcement learning on multi-hour tasks.
- Self-checking: improved self-verification and longer context management.
- Safety: a new safeguard stack; xAI reports a 3.3% risky-prompt pass-through rate on HackerBench v0.3.
- Distribution: live on day one in the Grok app, Cursor, Grok Build, the xAI API and GitHub Copilot.
How much does Grok 4.7 cost?
Grok 4.7 pricing is $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. The Fast variant doubles speed and doubles price, to $4 input and $12 output. That undercuts GPT-5.6 Sol by 50% on input and 70% on output, and Claude Fable 5.1 by 80% and 88%.
Those competitor list prices come from XenoSpectrum’s launch coverage: GPT-5.6 Sol at $4/$20 and Claude Fable 5.1 at $10/$50. xAI’s own line is that Grok 4.7 runs “twice as fast, at half the price of comparable models.”
To make the gap concrete, we priced a typical agent coding session: 10 million input tokens and 2 million output tokens. That ratio is common when an agent rereads a repository repeatedly.
| Model | Input / 1M | Output / 1M | Cost: 10M in + 2M out |
|---|---|---|---|
| Grok 4.7 | $2 | $6 | $32 |
| Grok 4.7 Fast | $4 | $12 | $64 |
| GPT-5.6 Sol | $4 | $20 | $80 |
| Claude Fable 5.1 | $10 | $50 | $200 |
At list price, one Fable 5.1 session buys about six Grok 4.7 sessions. For a team burning billions of tokens a month, that is not a rounding error. It is a line item a CFO will ask about.
How does Grok 4.7 score on benchmarks?
Grok 4.7 benchmarks look strong in xAI’s own table and weaker in independent testing. xAI reports 46.3% on CursorBench 4.0, 71.0% on DeepSWE v1.1 and 64.0% on EEBench. Artificial Analysis scores it 46 on its Intelligence Index v4.3.2, versus 53 for both Claude Fable 5.1 and GPT-6.
xAI’s numbers
On xAI’s launch page, CursorBench 4.0 rose to 46.3% from Grok 4.6’s 40.4%, a six-point jump. XenoSpectrum’s comparison table shows that edges GPT-5.6 Sol Max at 41.7% but trails Claude Fable 5.1 Max at 51.8%.
The best-looking result is EEBench, an electrical engineering test. Grok 4.7 posts 64.0% there, against 56.4% for Fable 5.1 Max and 39.4% for GPT-5.6 Sol Max. The SpaceX engineering data may be showing up exactly where you would expect it.
The independent numbers
Independent testing is less generous. The Decoder reports Artificial Analysis measured Grok 4.7 at 26% on Terminal-Bench 4.0. GPT-6 Astra scored 60% and Claude Fable 5.1 scored 55%. DeepSeek V4.1 Flash, an open-weights model, scored 27%.
Read that again. On agentic terminal work, xAI’s $2 flagship landed one point behind a Chinese open-weights model we covered in our DeepSeek V4.1 Flash breakdown. That comparison will not appear in xAI’s marketing.
| Benchmark | Grok 4.7 | Claude Fable 5.1 | GPT-6 / GPT-5.6 Sol | Source |
|---|---|---|---|---|
| AA Intelligence Index v4.3.2 | 46 | 53 | 53 (GPT-6) | Artificial Analysis |
| Terminal-Bench 4.0 (independent) | 26% | 55% | 60% (GPT-6 Astra) | Artificial Analysis |
| Terminal-Bench 4.0 (vendor) | 38% | 57.9% | — | xAI |
| CursorBench 4.0 | 46.3% | 51.8% | 41.7% (5.6 Sol Max) | xAI |
| DeepSWE v1.1 | 71.0% | 70% | 72.7% (5.6 Sol Max) | xAI |
| GDPval (Elo) | 1,695 | 1,735 | — | Decrypt |
The 12-point gap nobody explained
xAI’s own Terminal-Bench 4.0 figure is 38%. Artificial Analysis measured 26% on the same benchmark. That is a 12-point gap between vendor and independent runs, and xAI has not explained it.
Harness choice, effort settings and retries can all move agentic scores. But buyers should assume the independent number until xAI publishes its setup. Vendor benchmarks are marketing. Third-party runs are due diligence.
Is Grok 4.7 better than Claude Fable 5.1?
No, not on raw capability. Grok 4.7 trails Claude Fable 5.1 by seven points on the Artificial Analysis index and by 29 points on independent Terminal-Bench 4.0. It wins on price, on EEBench and, narrowly, on xAI-reported DeepSWE. For cost-sensitive, high-volume coding, it is a credible second-tier option.
Even Musk lowered the bar before launch. On September 14 he wrote: “Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others.” That is an honest framing, and it tells you where xAI itself places the model.
Decrypt also reports Grok 4.7 at 1,657 on AA-Briefcase, against 1,678 for Fable 5.1. On GDPval it sits 40 Elo points behind. Close on knowledge-work tasks. Far behind on long agentic loops.
Who wins and who loses from Grok 4.7?
Developer platforms and cost-conscious enterprises win. GitHub Copilot, Cursor and model routers get a cheap new option that pressures every Western lab’s pricing. The losers are Chinese mid-tier vendors whose main advantage was price, and anyone who hoped xAI would close the frontier gap this cycle.
The winners
GitHub rolled Grok 4.7 into Copilot on launch day across Pro, Pro+, Max, Business and Enterprise plans, according to XenoSpectrum. It runs in VS Code, JetBrains IDEs, Xcode, the Copilot CLI and cloud agents, billed pay-as-you-go at list price.
That default-on distribution matters more than any benchmark. New models are enabled automatically unless admins block them through model policy. xAI just got a free shelf in front of millions of developers.
The losers
The Decoder describes Grok 4.7’s pricing as closer to Chinese models than to Western frontier alternatives. That squeezes vendors like StepFun, whose Step 5 Preview launched at $1 input a day earlier. The Western premium for “not Chinese” is now available at near-Chinese prices.
It also complicates the math for coding startups building on cheap bases. Cognition’s approach, covered in our SWE-2 analysis, relied on a Chinese open model to cut costs. A US-hosted $2 option changes the compliance calculus for enterprise buyers.
Why did Grok 4.7 take so long?
Grok 4.7 slipped at least five times. Musk promised it “in 4 weeks” on July 24, then “3 to 4 weeks” on August 12, then “10 days” on September 1. Reports tied the delays to reinforcement-learning and self-checking problems. It shipped September 21, about a month after the first target.
CellCog’s timeline shows Musk saying on September 11 that the model “needs a few more days to cook.” Each slip mattered because rivals kept shipping. GPT-6 Astra and Claude Fable 5.1 both landed during the delay window, moving the target xAI was chasing.
xAI is already selling the next one. Decrypt reports the roadmap as Grok 4.8, a “meaningful step up,” then Grok 4.9 at “Astra/Fable class,” then Grok 5. None has a date. Given this launch’s record, investors should discount those promises accordingly.
What does Grok 4.7 mean for AI pricing?
Grok 4.7 confirms a two-tier market. A small frontier tier, led by Claude Fable 5.1 and GPT-6, charges premium prices for the hardest agentic work. Below it, a crowded “good enough” tier is racing toward $1 to $2 per million input tokens. Margin lives at the top.
That is the real signal for capital. xAI kept Grok 4.6’s price while raising capability, which means it is buying share, not harvesting margin. Anthropic and OpenAI can hold premium pricing only while their independent benchmark lead holds. Today that lead is seven index points.
xAI’s broader legal and competitive fights, including the Apple antitrust case it dropped last week, show a company that wants distribution badly. Copilot on day one delivers some of it.
Frequently asked questions
When was Grok 4.7 released?
xAI released Grok 4.7 on September 21, 2026, after at least five delays since late July.
How much does the Grok 4.7 API cost?
$2 per million input tokens and $6 per million output tokens. The Fast variant costs $4 and $12.
Is Grok 4.7 available in GitHub Copilot?
Yes. It rolled out on launch day to all Copilot plans, including Pro, Business and Enterprise, billed at list price.
Is Grok 4.7 better than GPT-6?
No. Artificial Analysis scores GPT-6 at 53 and Grok 4.7 at 46. On Terminal-Bench 4.0, GPT-6 Astra scored 60% versus 26%.
How many parameters does Grok 4.7 have?
Elon Musk has said 2.1 trillion. xAI has not published an official specification.
Was Grok 4.7 trained on SpaceX data?
Musk said SpaceX’s non-ITAR engineering corpus was used in supplemental training. xAI has not detailed the dataset.
The bottom line
Grok 4.7 is a price weapon, not a frontier model. At $32 for a workload that costs $200 on Claude Fable 5.1, it will win volume in Copilot and Cursor. But a 12-point gap between xAI’s and independent Terminal-Bench scores is a red flag.
Verdict: use it for high-volume, lower-stakes coding and engineering tasks. Keep Fable 5.1 or GPT-6 for long autonomous agents. And wait for independent numbers before believing any Grok 4.9 promise.
Leave a Reply