OpenAI's GPT-5.6 Sol Just Got 14 Times Faster: Here's Why Speed Matters Now
OpenAI introduced Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing, generating up to 750 output tokens per second. The company announced the preview service on August 13, initially rolling it out through its application programming interface (API) to a limited group of customers, with broader access planned as computing capacity grows.
What Makes This Speed Breakthrough Different?
The Ultrafast mode represents a fundamental shift in how large language models (LLMs) can be deployed in real-world workflows. Output tokens are pieces of generated text, and the ability to produce 750 per second means the model can respond to complex requests almost instantly. Traditionally, users had to choose between a smaller, faster model or a more capable one that took longer to respond. Ultrafast aims to eliminate that trade-off by preserving GPT-5.6 Sol's full capabilities while cutting response time dramatically.
The speed improvement is powered by Cerebras, a chipmaker specializing in low-latency artificial intelligence (AI) compute infrastructure. OpenAI announced its partnership with Cerebras in January, adding 750 megawatts of computing power to its platform. This infrastructure foundation enabled the company to build Ultrafast as a dedicated service tier.
Which Industries and Workflows Benefit Most?
OpenAI identified several use cases where the speed advantage creates immediate value. The company expects the most impact in incident response, financial research, customer support, commerce, and live research applications. During an outage, engineers can now review logs, traces, and recent code changes while the model responds in near real-time, rather than waiting for batch processing. Support systems can handle multi-step customer requests during live conversations without noticeable delays.
Early customers have already reported transformative results. John Crepezski of Jane Street, a quantitative trading firm, called the speed increase "impressive" and noted that it makes focused work alongside models more practical. For voice AI applications, the speed advantage is particularly significant.
"The speed completely changes the call experience on complex work," said Courtland Lykins, product lead for voice AI at Podium.
Courtland Lykins, Product Lead for Voice AI at Podium
Financial researchers represent another key beneficiary. They can now analyze changing market signals and transactions without waiting for slower batch-style processing. This capability is especially valuable when market conditions shift rapidly and delayed analysis becomes outdated.
How to Evaluate Whether Ultrafast Fits Your Workflow
- Real-Time Responsiveness: If your application requires immediate answers to complex questions, such as customer support chatbots or live trading analysis, Ultrafast's 14-fold speed improvement could eliminate noticeable delays that frustrate users.
- Interactive Work Patterns: Applications where humans work alongside the model in real-time, like code review during outages or voice conversations, benefit more than batch processing systems that can tolerate longer wait times.
- Infrastructure Constraints: The service remains in limited preview, constrained by available computing capacity rather than technical limitations. Access expands as OpenAI builds more infrastructure, so early adoption depends on being selected for the preview program.
OpenAI emphasized that the preview remains constrained by infrastructure rather than a broad public rollout. The company is using feedback from initial customers to determine where the speed produces the most value before expanding availability to a wider audience.
This development builds on OpenAI's earlier work with Cerebras. In February, GPT-5.3-Codex-Spark became the first model tied to the partnership, using Cerebras hardware to exceed 1,000 tokens per second in real-time coding applications. Ultrafast extends that approach to the company's flagship GPT-5.6 Sol model, making high-speed inference available for a broader range of use cases.
The timing of Ultrafast's launch reflects a broader industry shift toward optimizing AI models not just for accuracy, but for speed and cost-efficiency in production environments. As AI becomes embedded in customer-facing applications, the ability to deliver responses in milliseconds rather than seconds becomes a competitive advantage. OpenAI's infrastructure partnership with Cerebras appears designed to sustain that advantage as demand grows.