Back
LLMs

Grok 4.7 Explained: Specs, Price, and How It Stacks Up

6 min read

Abstract 3D rendering of a glowing sphere made of connected dots and lines, representing an AI neural network

Photo by Growtika on Unsplash

Grok 4.7 Explained: Specs, Price, and How It Stacks Up

SpaceXAI (the company formerly known as xAI, now folded into SpaceX after their July 2026 merger) shipped Grok 4.7 on September 21. It's a bigger model than Grok 4.6, trained for longer with reinforcement learning, and aimed squarely at coding and agentic work. The pricing didn't move: still $2 per million input tokens and $6 per million output tokens, the same rate xAI has held since the previous release.

That combination, a larger model at flat pricing, is the headline. But the benchmark numbers that came out alongside the launch tell a more complicated story about where Grok actually sits in the current model lineup.

What actually changed in Grok 4.7

According to SpaceXAI's own announcement, three things separate 4.7 from its predecessor. First, it's a larger base model, roughly 40% more parameters based on early estimates circulating on Hacker News and elsewhere, though the company hasn't confirmed an exact figure. Second, the reinforcement learning phase ran longer and on harder, multi-hour tasks rather than short single-turn problems. Third, the model was built to check its own work before returning an answer, what SpaceXAI calls improved self-verification.

There's also a new safety layer. The company describes an "entirely new safeguard stack" with tighter jailbreak resistance, and points to a benchmark called HackerBench where only 3.3% of risky dual-use prompts got through. Whether that number holds up under independent red-teaming is a separate question, but it's a deliberate response to criticism SpaceXAI has faced over Grok's guardrails in past releases.

On the product side, Grok 4.7 ships in the Grok API, Grok Build, and Cursor, and it landed in GitHub Copilot as a selectable model the same week. A faster variant is available too: double the price for double the output speed, aimed at latency-sensitive agent workflows rather than long reasoning chains.

The benchmark picture

This is where the launch gets less flattering. Independent scoring from Artificial Analysis puts Grok 4.7 at 46 on its Intelligence Index (v4.3.2), well behind both Claude Fable 5.1 and GPT-6, which each scored 53. The gap widens on agentic coding specifically: Grok 4.7 managed 26% on Terminal-Bench 4.0, compared to 60% for GPT-6 Astra and 55% for Claude Fable 5.1.

SpaceXAI's own release notes lean on a different set of benchmarks, including CursorBench 4.0 (46.3%), DeepSWE v1.1 at high effort (71.0%), and domain-specific tests like HealthBench Professional (56.7%) and the Harvey Legal Agent Benchmark (19.6%). Worth noting: these are vendor-run numbers, so treat them as a starting point rather than a verdict. The Artificial Analysis figures, which use a standardized methodology across vendors, are the more useful comparison if you're trying to rank Grok against competitors rather than just see how it did on its own scale.

Put plainly: Grok 4.7 is not catching up to the current frontier on general intelligence or coding agent benchmarks. What it's competing on is price.

Why the pricing matters more than the benchmarks

At $2/$6 per million tokens, Grok 4.7 undercuts most Western frontier models by a wide margin and sits closer to what Chinese labs have been charging. For teams running high-volume workloads, batch document processing, customer support triage, first-pass code review, that price gap can matter more than a 10-point difference on an intelligence index, especially for tasks where a 46 is good enough.

This is really the pattern SpaceXAI has followed since Grok 4.0: don't chase the top of the leaderboard, compete on throughput and cost instead. It's a coherent strategy, but it does mean Grok 4.7 isn't the model to reach for if you need the strongest possible reasoning or the highest agentic coding success rate. It's the model to reach for when you're running the same task a million times and the cost curve actually matters.

Where it fits against GPT-6 and Claude Fable 5.1

If you're picking a model for a new project this week, the rough breakdown looks like this: GPT-6 Astra and Claude Fable 5.1 lead on raw coding agent performance and general reasoning. Grok 4.7 is meaningfully cheaper and fast enough for latency-sensitive pipelines, but you'll want to test it against your specific task before committing, especially anything that resembles multi-step agentic coding, where the Terminal-Bench gap is largest.

For general knowledge work, customer-facing chat, and simpler coding assistance, the gap narrows enough that price and integration (Cursor, Copilot, existing API tooling) may decide it for you.

Should you switch to Grok 4.7?

If you're already building on Grok 4.6, upgrading is close to a no-brainer since the price hasn't changed and the new model outperforms the old one on nearly every benchmark SpaceXAI published. The harder question is whether to migrate from GPT-6 or Claude Fable 5.1, and the honest answer depends on what your workload actually looks like rather than what the marketing page says.

For anything approaching autonomous multi-step coding, opening pull requests, running test suites, chasing down build failures across a codebase, the Terminal-Bench gap is large enough that it's worth running your own eval before switching, not just trusting the vendor benchmarks. A 26% success rate versus 55-60% isn't a rounding error. But if your use case is closer to single-turn code generation, documentation, summarization, or high-volume classification, the practical difference between a 46 and a 53 on a general intelligence index is much smaller than the difference in your monthly bill.

It's also worth watching how the safety stack performs once it's been in the wild for a few weeks. SpaceXAI's own HackerBench numbers look strong, but those are self-reported, and Grok's past releases have drawn scrutiny over content moderation and jailbreak resistance. Independent testing usually surfaces issues that internal benchmarks miss, so treat the 3.3% figure as a starting claim rather than a settled fact.

Key takeaways

Grok 4.7 launched September 21 with a larger base model, longer reinforcement learning, and the same $2/$6 pricing as Grok 4.6. Independent benchmarks from Artificial Analysis place it behind GPT-6 and Claude Fable 5.1 on both general intelligence and agentic coding, with the largest gap on Terminal-Bench 4.0. Its real selling point is cost and speed, not top-end capability. It's already available in the Grok API, Cursor, Grok Build, and GitHub Copilot.

FAQ

Is Grok 4.7 better than GPT-6? Not on the benchmarks that have been independently verified so far. GPT-6 Astra scores noticeably higher on both general intelligence and agentic coding tests. Grok 4.7's advantage is price, not raw capability.

How much does Grok 4.7 cost? $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.6. A faster variant costs twice as much for roughly double the output speed.

Where can I use Grok 4.7? It's live in the Grok API, Grok Build, Cursor, and as a selectable model in GitHub Copilot, plus assorted third-party coding platforms and model routers.

Is the "SpaceXAI" name a typo for xAI? No. xAI merged into SpaceX and adopted the SpaceXAI branding in July 2026, so recent official material uses that name.

  • Grok 4.7
  • SpaceXAI
  • LLM benchmarks
  • AI coding tools
  • GPT-6