SpaceXAI launches Grok 4.6 coding flagship, 500K context at $2/$6 per million tokens
On Aug 13 SpaceXAI (formerly xAI) released Grok 4.6, targeting long-horizon agents and coding with 500K context and four reasoning tiers.
On August 13, SpaceXAI officially released Grok 4.6, its new flagship AI model, further shifting its upgrade focus toward long-horizon agent scenarios compared to Grok 4.5. Grok 4.6 targets complex tasks requiring continuous planning, execution, and iteration, applicable to large codebases, bug fixing, web research, data analysis, and web application building.
The model supports text and image inputs, with an API context window of 500K tokens and capabilities including Function Calling and Structured Outputs. SpaceXAI trained the model on agent environments spanning general programming, software engineering, web development, kernel optimization, and CAD, focusing on improving long-chain goal retention, tool invocation, and self-verification.
On the Artificial Analysis Intelligence Index (a nine-benchmark composite), Grok 4.6 scored 61, tying GPT-5.6 Sol Max and trailing Claude Fable 5 Max (62) by one point. Engineering benchmarks are particularly strong: DeepSWE v1.1 reached 65.9% (+11.9 points vs Grok 4.5), CursorBench v3.2 hit 69.9%, and Terminal-Bench v3.0 doubled to 26%. On the AA-Briefcase long-task benchmark, Grok 4.6 scored 1577, above GPT-5.6 Sol Max (1502) and Fable 5 Max (1574).
Standard API pricing is $2 per million input tokens and $6 per million output tokens, unchanged from Grok 4.5. Long-context usage above 200K prompt tokens is billed at $4 input and $12 output per million, with cached input rising from $0.50 to $1 at the same threshold. Grok 4.6 is currently available via Grok products, the SpaceXAI API, and third-party platforms including Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare. Model weights are not open-sourced.