Grok 4.7 undercuts frontier rivals at half the token cost
Grok 4.7 delivers 46.3% on CursorBench 4.0—a 6-point jump—at $2/$6 per million tokens. That's ~87% of frontier-class performance at roughly half the cost. An average task runs $6.01 vs GPT-5.6 Sol's $8.23, a 27% cut. New 500K-token context and a 96.7% dual-use blocking rate target high-volume enterprise coding. The catch: it burns 81K output tokens per task, up from 36K. Cheaper per job, hungrier per token—can that trade-off flip enterprise contracts?
On September 21, 2026, xAI quietly replaced Grok 4.6 as its default coding agent model. The upgrade landed with a familiar price tag—$2 per million input tokens and $6 per million output tokens—yet it signals something more consequential than a routine refresh. Grok 4.7 represents a deliberate strategy: compete on capability first, let the price do the marketing.
The Spec Sheet That Matters
The headline feature is a 500,000-token context window, up sharply from its predecessor and now comfortably in the range of top-tier rivals like Claude Fable 5.1 and GPT-6. The model was built on a ~2.1-trillion-parameter base—40% larger than Grok 4.6's 1.5T—then subjected to a longer reinforcement-learning run focused on harder, multi-hour tasks, with SpaceX engineering datasets (Starlink telemetry, manufacturing logs, failure analysis) folded into training. That emphasis shows up where it counts: self-verification of its own work, long-context management, and reasoning across four effort levels—low, medium, high, and a new "xhigh" tier (with high as the default).
The result is measurable, not just claimed. On the CursorBench 4.0 multi-hour coding benchmark, Grok 4.7 scores 46.3%, up from 40.4% for Grok 4.6 and ahead of GPT-5.6 Sol's 41.7%. On DeepSWE v1.1 at xhigh effort, it reaches 71.0%. It completes the full coding agent loop—finding files, implementing changes, running tests—plus research assistance with citation tracking and document creation with editable output. The gains extend beyond coding: Terminal-Bench 4.0 jumped from 20.3% to 38.0%, a 17.7-point swing that underscores the RL run's focus on shippable, multi-hour work.
Where It Wins, and Where It Doesn't
The competitive picture is more nuanced than the release hype suggests. On the widely cited Artificial Analysis index, Grok 4.7 scores 46, compared with 53 for both Claude Fable 5.1 and GPT-6—a two-point jump over Grok 4.6's 44 that nonetheless places it in the second tier on raw power. But on price-per-performance grounds, the calculus flips: at roughly half the token cost of leading Western models, Grok 4.7 delivers roughly 87% of their headline index score. For high-volume coding workloads that run into the millions of tokens, that ratio changes the economics of deployment—an average task cost of $6.01 versus $8.23 for GPT-5.6 Sol, a ~27% reduction.
Safety is a differentiator xAI leaned into explicitly. The model ships a redesigned safeguard stack, and the numbers argue for genuine dual-use protection: a 62.4% LatchBio score and a 3.3% risky-prompt allowance on HackerBench, blocking 96.7% of dangerous dual-use prompts. Those figures indicate the model resists most attempts to elicit dangerous biological or cyber content—relevant given its visibility in enterprise tooling.
The Distribution Play
Grok 4.7's reach extends beyond xAI's own API. It's immediately available in Cursor's model picker, Grok Build, third-party coding harnesses, model routers, cloud platforms, and GitHub Copilot across every subscription plan—a wide surface that pushes the model into high-volume developer workflows, but with an efficiency caveat: it consumes roughly 81,000 output tokens per task versus about 36,000 for Grok 4.6, a trade-off to monitor. The token appetite is offset at checkout—per-task cost lands at $4.69 versus $5.20 for Grok 4.6. A faster variant is exclusive to Cursor and Grok Build, tiering that pushes users into xAI's partner ecosystems. The company also signals expansion into the UAE market, including integration via Tesla vehicles.
The Outlook
The trajectory points to sustained incremental gains: longer RL runs, larger bases, broader integration. The benchmark record is not uniform—regressions on AA-Omniscience accuracy (-1 point) and AutomationBench-AA (-1.1 points) suggest the RL run traded away some breadth for depth. By holding pricing flat while narrowing the capability gap, xAI is selling the frontier at a discount—and betting that volume, distribution, and safety credentials can flip the next round of enterprise contracts. In a market where the crown is decided as much by deployment economics as by benchmark scores, that is a counteroffensive worth tracking.
Comments ()