xAI releases Grok 4.6
xAI released Grok 4.6, a closed model it says builds on Grok 4.5 for long-running agents and interactive visual work. On the Artificial Analysis Intelligence Index, a composite of nine benchmarks, it scores 61, tying GPT-5.6 Sol Max and sitting just behind Fable 5 Max at 62, and it is live in Cursor, Grok Build, and the API at $2 per million input tokens and $6 per million output tokens.
xAI released Grok 4.6 on August 12, 2026. The company describes it as a successor to Grok 4.5 built for long-running agents and more ambitious interactive and visual work: staying with a complex task across many steps, whether that is researching a topic, working across a codebase, or turning an idea into a polished application. Weights are closed.
Some background makes the release more interesting than the version bump suggests. Grok 4.5 launched on July 9 with unusual gaps. xAI called it an Opus-class model running on its 1.5-trillion-parameter V9 foundation, but published no price, no context window, and no results on any public leaderboard. Grok 4.6 fills in part of that picture.
How it compares to Grok 4.5
xAI now cites the Artificial Analysis Intelligence Index, a composite score built from nine separate benchmarks. Grok 4.6 lands at 61. That ties OpenAI's GPT-5.6 Sol Max, also at 61, and sits one point behind Anthropic's Fable 5 Max at 62. Grok 4.5 High scores 56 on the same index, so the generational jump is five points.
Three of the underlying tests are worth unpacking. CursorBench 3.2, which measures real coding work inside an editor, has Grok 4.6 at 69.9%, up from 66.7% for Grok 4.5 High and ahead of GPT-5.6 Sol Max at 67.2%. Fable 5 Max still leads at 70.5%. DeepSWE 1.1, a broader software engineering test, shows the biggest improvement: 65.9% against 54% for Grok 4.5 High, though GPT-5.6 Sol Max at 73% and Fable 5 Max at 70% remain ahead. And Terminal-Bench 3.0, which tests work in a real command-line environment, is the weak row. Grok 4.6 scores 26%, well up from 15.7% but far behind both rivals in the mid-30s.
All of these are xAI's numbers. Competitor figures, the company says, are drawn from developers' system cards or public leaderboards, taking the best of self-reported or published results. Until someone independent runs their own evaluation, the whole table is a company claim. xAI also says it saw the model do more self-testing on longer jobs, checking its own work before moving on, which is a vendor observation of the same kind.
What it costs and where to get it
Pricing is $2 per million input tokens and $6 per million output tokens. There is no earlier price to compare against, because xAI never published one for Grok 4.5; this is the first sticker price on the Grok line. A token is roughly three quarters of a word, and output costs three times the input rate. A faster variant runs at double the base price.
The model is live in Cursor, Grok Build, and the xAI API, and is also listed on OpenRouter, Vercel, and Cloudflare. Cursor and Grok Build subscribers need no new plan, and both get double their included usage for the first week.
What is still unknown
The context window, the amount of material the model can consider at once, is unpublished, as it was for Grok 4.5. Whether the fast variant gives up any quality is unstated. And on safety, xAI says safeguards were improved and calibrated to the model's capabilities, with its widest-ever pre-deployment testing suite plus post-deployment and third-party testing, but it names no testers and publishes no safety results.
Sources
- xAI x.ai
LMTimeline writes its own account of each event. Primary sources are linked above so you can read them directly.