OpenAI Tests Ultrafast Inference Mode, GPT-5.6 Sol Output Reaches Up to 750 Tokens per Second
2026-08-14 10:15
Favorite

en.Wedoany.com Reported - On August 13, OpenAI began previewing a high-speed inference mode called Ultrafast to a small group of customers, designed to accelerate the output speed of its flagship model GPT-5.6 Sol. The mode is currently in limited testing and is not yet available to all ChatGPT or API users.

The report, citing OpenAI, states that Ultrafast can reach processing speeds up to 14 times faster than the standard mode, with an output throughput of up to 750 tokens per second. Its primary goal is to improve the response speed of complex models in real-time business scenarios without switching to smaller models.

It is important to note that 750 tokens per second primarily reflects the output throughput of the model when generating text, and does not mean that all tasks—including inference, tool calls, file processing, and web retrieval—can be reduced to one-fourteenth of their original time. The end-to-end speed of different tasks will still be affected by inference intensity, context length, and tool execution time.

TechCrunch reports that the Ultrafast preview is supported by a collaboration between OpenAI and AI chip company Cerebras, with initial applications targeting enterprise workflows sensitive to response latency, such as security incident response, customer service, financial market analysis, and e-commerce. OpenAI plans to gradually expand access as computing capacity increases.

According to OpenAI's official model documentation, GPT-5.6 Sol is the flagship model of the GPT-5.6 series, designed for complex coding, research, computer operations, and professional work. The model supports a 1.05 million-token context window, up to 128,000 tokens of output, and multi-level inference intensity ranging from none to max. Ultrafast is not a standalone new model, but rather an operating mode for the inference service speed of GPT-5.6 Sol.

As of August 14, the GPT-5.6 Sol model page on OpenAI Docs and the official product update summaries from August 10 to 14 have not yet listed Ultrafast's pricing, API invocation methods, service tiers, regional availability, or the specific list of users granted access.

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com