Grok 4.7 Thinks Longer for the Same Money, and Wins an Odd Benchmark
xAI shipped Grok 4.7 on Monday at unchanged prices of 2 dollars in and 6 dollars out per million tokens. It posts 71 percent on DeepSWE and an unusually high score on a legal benchmark where rivals collapse.
xAI released Grok 4.7 on Monday, and the headline is what did not change: it costs the same as Grok 4.6, at 2 dollars per million input tokens and 6 dollars per million output tokens, with a faster variant at double the price for double the speed. A token is roughly three quarters of a word, so a million of them is a long afternoon of work for a coding assistant. Under the hood it is a bigger base model trained with a longer reinforcement learning run, weighted toward problems that take hours rather than minutes.
The benchmark table is a mixed picture, which is more useful than a clean sweep. On DeepSWE v1.1, a software engineering test, Grok 4.7 gets 71.0 percent at high effort, up from 65.2 for its predecessor and just ahead of Claude Fable 5.1 Max at 70.0, while GPT-5.6 Sol Max leads at 72.7. On Terminal-Bench 4.0, which measures multi-hour command line work, Grok 4.7 nearly doubles its predecessor’s score to 38.0 percent, yet Fable 5.1 Max is far ahead at 57.9. The oddest number is the Harvey Legal Agent Benchmark, where Grok 4.7 posts 19.6 percent against 2.5 for GPT-5.6 Sol and 6.7 for Fable 5.1 Max. On clinical reasoning it trails both. xAI also says this is its strongest model yet at refusing genuinely dangerous requests while still helping with legitimate security work, letting through 3.3 percent of risky prompts on its own HackerBench test.
What is behind this
Two things are worth pulling out. First, prices are not falling so much as capability is rising underneath a fixed price, which is the pattern across the industry this year and the reason a model from six months ago starts feeling expensive rather than getting cheaper. Second, the spread across those benchmarks tells you something real: there is no single “best model” any more, there are models that suit particular jobs. A score of 19.6 percent still means four out of five legal tasks failed, so the right reading is that these agents are early at professional work, not that Grok has solved law. Benchmarks are also run by the vendor unless stated otherwise, and each lab picks the tests where it looks good.
What this means for you: For everyday chat use, this changes very little. If you write code or pay for AI by the token, it is worth a trial, particularly because it sits inside tools you may already use such as Cursor and the Grok API, and because nothing about your bill changes when you switch. The more general lesson: when a model gets better at the same price, your existing budget quietly buys more, but only if you actually move. Check every few months which model your tools default to.
Sources
Source: https://x.ai/news/grok-4-7
When Agents Got Chatty, This Startup's Margins Went to Minus 50 Percent
Bloomberg reports legal AI firm Harvey watched gross margins fall from about 50 percent to minus 50 by June as token use rose twentyfold. It recovered by building an in-house model on Moonshot's open Kimi K3.