SpaceXAI (formerly xAI) releases Grok 4.6, a new flagship AI model built to handle long-running, agentic tasks such as researching topics, working through codebases, and turning product ideas into working applications. The company makes the model available the same day through its API, its Grok Build tool, and the Cursor code editor, with additional distribution via partners such as OpenRouter, Vercel, and Cloudflare. SpaceXAI says it also improves multi-step task reliability, including more self-testing and verification during longer trajectories.

The launch follows Grok 4.5 and reflects a longer supplemental training effort, including curated model-generated reasoning data, changes to the optimizer and training recipe, regenerated supervised fine-tuning trajectories, and reinforcement learning across agentic environments. On benchmarks compiled from third-party and reported results, Grok 4.6 matches OpenAI’s GPT-5.6 Sol on the Artificial Analysis Intelligence Index at a score of 61, with additional reported gains in coding and knowledge-work tests versus Grok 4.5. VentureBeat also emphasizes that pricing and efficiency claims depend on how many turns and tokens a workflow uses, noting Grok 4.6’s API context window and different billing rates for large prompts.

Outlets also differ in emphasis: VentureBeat highlights enterprise cost-per-task comparisons and broader risks related to the Grok brand and prior governance controversies, while Free Press Journal focuses more on the model’s agentic capabilities, training approach, rollout, and pricing terms. Both sources cite the same core availability and $2 per million input tokens / $6 per million output tokens starting API pricing.