en.Wedoany.com Reported - On August 12, U.S. AI company SpaceXAI released its new flagship model, Grok 4.6, with a focus on enhancing long-horizon agentic tasks, software development, knowledge work, and interactive application building capabilities. The model is now available via Grok Build, Cursor, and the SpaceXAI API, and has been integrated into partner platforms including OpenRouter, Vercel, and Cloudflare.

Grok 4.6 supports text and image input with text output, featuring a 500K token context window. It can invoke functions, web search, X platform search, and code execution tools, and supports structured output. Developers can access the model through the Responses API and Chat Completions interface, with reasoning effort adjustable across four levels—low, medium, high, and ultra-high—with high reasoning mode set as the default.
SpaceXAI stated that Grok 4.6 is primarily designed for complex tasks requiring the sequential execution of multiple steps, including cross-source research, information analysis, large codebase processing, software prototyping, and work product generation. Compared with Grok 4.5, the new model performs more self-testing and result verification during long tasks, and strengthens its ability to generate initial versions of interactive applications from product concepts.
In terms of training, Grok 4.6 employs a longer supplemental training process than the previous generation, with data including model-generated data for reasoning and advanced technical concepts, as well as engineering data. SpaceXAI subsequently used Grok 4.5 to regenerate supervised fine-tuning trajectories covering different reasoning effort levels, agentic frameworks, science, engineering, mathematics, and software development domains, and filtered problematic training records through model checks. The model also underwent reinforcement learning training on agentic tasks such as knowledge work, general programming, kernel optimization, web development, and computer-aided design.
According to data released by SpaceXAI, Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index composite benchmark, an improvement of 5 points over Grok 4.5's 56, tying with GPT-5.6 Sol and trailing Claude Fable 5 by 1 point. This index aggregates nine tests spanning science, programming, and financial services, but the scores primarily reflect model performance under specific conditions and cannot be directly equated with capabilities across all real-world business scenarios.
In agentic and software development benchmarks, Grok 4.6 achieved 69.9% on CursorBench 3.2, 65.9% on DeepSWE 1.1, and 61.3% on the FrontierCode 1.1 extended test. Its APEX-Agents score was 57.5%, up from Grok 4.5's 47.1%; on the AA-Briefcase test, which evaluates complex knowledge work capabilities, it scored 1577, also higher than the previous generation's 1313.
The standard Grok 4.6 API base pricing is $2 per million input tokens, $0.5 per million cached input tokens, and $6 per million output tokens. When input context exceeds 200K tokens, the official rate increases; SpaceXAI also offers a faster version priced at approximately twice the standard rate. Grok Build and Cursor provide double usage credits during the first week following the model's release.
SpaceXAI stated that the new model has completed the company's most extensive pre-deployment capability and safety testing to date, with continued post-deployment and third-party testing planned. However, most of the officially released performance data comes from the company's own testing or cites publicly available results from other model developers, and still requires further validation through independent evaluations and real-world enterprise deployments.





















