OpenAI has introduced a new operating mode for its flagship model that dramatically increases processing speed, addressing one of the most common requests from users who rely on the technology for time-sensitive tasks.
The San Francisco-based AI company has officially launched what it calls Ultrafast, a mode built specifically to accelerate the output of its most capable model, GPT-5.6 Sol. According to OpenAI, the mode can operate at up to 14 times the speed of standard processing, generating as many as 750 output tokens per second. Tokens, in this context, refer to the discrete units of text that a large language model produces during a conversation or task.
A New Direction for Speed and Capability
In a blog post published Thursday, OpenAI framed the development as a meaningful shift in how the industry thinks about the relationship between speed and model quality. Historically, achieving near-real-time response speeds required developers to opt for smaller, more narrowly focused models rather than the most powerful ones available.
"Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second."
The statement signals OpenAI's intent to close the gap between performance and speed - a trade-off that has long shaped how businesses and developers choose which AI models to deploy in production environments.
Competitive Landscape
OpenAI is not alone in pursuing faster inference speeds. Rival AI company Anthropic has also introduced an accelerated tier for its Claude model family. However, the speeds being claimed by OpenAI with Ultrafast appear to exceed what Anthropic currently offers through its own fast mode, positioning this release as a notable step forward in the broader competitive race among AI providers.
Enterprise Applications
OpenAI has outlined several business use cases where it believes Ultrafast could offer the most immediate value. The company specifically highlighted the following areas as strong candidates for deployment:
- Incident response and IT operations management
- Customer service and technical support workflows
- Financial market analysis and real-time data processing
- E-commerce applications requiring rapid, dynamic responses
These use cases share a common requirement - the need for fast, accurate, and contextually relevant outputs under time pressure. By combining the reasoning capabilities of GPT-5.6 Sol with dramatically accelerated processing, OpenAI is positioning Ultrafast as a tool suited for enterprise environments where delays carry tangible costs.
Powered by Cerebras Hardware
The technical backbone behind Ultrafast is OpenAI's existing partnership with Cerebras, a semiconductor company known for developing chips specifically designed to handle the computational demands of large-scale AI workloads. The collaboration enables the kind of throughput that makes 750 tokens per second achievable at scale.
Ultrafast is currently available in a limited preview, with access restricted to a select group of customers. OpenAI confirmed that it plans to broaden availability over time as its infrastructure capacity expands, suggesting that wider rollout will be tied directly to how quickly the company can scale the underlying hardware resources supporting the feature.
The release reflects a broader trend across the AI industry, where companies are increasingly investing not just in making models smarter, but in making them faster and more practical for real-world deployment at enterprise scale.



