Tech

OpenAI Unveils ‘Ultrafast’ Mode for GPT 5.6 Sol in Partnership with Cerebras

The AI lab’s latest offering targets high-volume corporate tasks including financial market analysis and incident response, though access remains restricted to a select group of customers for now.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · View original source
OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT 5.6 Sol work at 14x the speed
New processing tier promises 14x speed increase and up to 750 tokens per second for enterprise workflows

OpenAI has introduced a preview of ‘Ultrafast’, a new processing mode for its GPT 5.6 Sol model that increases speed by 14 times, delivering up to 750 output tokens per second. The feature, powered by a partnership with chipmaker Cerebras, targets enterprise workflows including incident response, customer service, and financial market analysis. Access is currently restricted to a small group of customers.

The launch marks a strategic shift in how the company approaches latency, moving away from the traditional trade-off between model capability and speed. In a blog post published on Thursday, OpenAI noted that achieving real-time performance previously required selecting smaller or more specialised models. The new Ultrafast mode aims to deliver more useful work per second without compromising the utility of the GPT 5.6 Sol, which is described as the company’s latest and most powerful model.

Output tokens, which represent the distinct pieces of text generated by a large language model during interaction, are the primary metric for this performance gain. By leveraging Cerebras hardware, OpenAI claims the new mode can process data at rates previously unattainable with its flagship architecture. The company states that performance quality remains uncompromised despite the accelerated processing speeds.

The technology is being rolled out initially to a limited audience as part of a preview phase. OpenAI has indicated that it plans to expand access to the feature as infrastructure capacity grows, though no specific timeline has been provided for broader availability. This cautious rollout allows the firm to monitor system stability and performance metrics before scaling the service to a wider enterprise base.

OpenAI’s move positions it competitively against rivals such as Anthropic, which offers a ‘fast mode’ for its Claude models. However, OpenAI claims its new offering is superior in speed, noting that Anthropic’s accelerated option does not deliver comparable throughput. The focus on high-speed processing reflects broader industry trends where latency is becoming a critical factor for real-time corporate applications.

Primary use cases for Ultrafast include incident response, customer service, financial market analysis, and e-commerce. These sectors require rapid data processing and immediate feedback loops, making the 14x speed increase particularly relevant for time-sensitive decision-making processes. The availability of such speed on a powerful model like GPT 5.6 Sol could influence how enterprises integrate generative AI into their operational workflows.

While the immediate impact on performance quality at 14x speed is not fully detailed beyond the company’s claims, the introduction of Ultrafast highlights the ongoing race to optimise large language model efficiency. As capacity expands, the feature could become a standard component for enterprises requiring both high intelligence and low latency in their AI deployments.

Continue reading

More from Tech

Read next: Namecheap services disrupted by Phoenix data centre cooling failure
Read next: Trump administration authorises private firms to conduct cyber attacks on overseas criminals
Read next: Virgin Galactic opens public vote to name new Delta-class spaceship