On August 13,

This feature was jointly developed by OpenAI and chip manufacturer Cerebras, currently in preview phase, and is initially available to selected enterprise customers. It will gradually expand its coverage as computing power increases.
Traditional large model applications have long been constrained by the trade-off between "model capability" and "response speed," often forcing real-time scenarios to use smaller parameter models. The implementation of the Ultrafast mode breaks this dilemma, enabling GPT-5.6Sol to be seamlessly deployed in complex workflows such as event response, customer service, financial market analysis, and e-commerce, which are extremely sensitive to latency.
Compared to the fast mode of its competitor Anthropic's Claude, OpenAI demonstrates stronger advantages in output throughput and execution efficiency. This technological breakthrough marks the acceleration of cutting-edge large models from "high-latency offline inference" toward "high-throughput real-time interaction," and will reshape the deployment form and productivity standards of enterprise-level AI in the future.
