Tech

Cerebras and OpenAI Launch Ultrafast Mode for GPT-5.6 Sol

Benchmarks show significant speed advantages over competitors, with Cerebras hardware eliminating data movement bottlenecks to accelerate critical AI workflows.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
New service tier delivers frontier intelligence at up to 750 output tokens per second

Cerebras and OpenAI have announced the launch of Ultrafast Mode, a new service tier within the OpenAI API powered by Cerebras hardware. The mode delivers the GPT-5.6 Sol model at speeds of up to 750 output tokens per second, with claims that performance quality remains uncompromised. Initial access is restricted to a select group of customers, with broader availability planned as capacity expands.

The development addresses a longstanding trade-off in artificial intelligence, where larger, more intelligent models typically incur higher computational costs and slower response times. By utilising Cerebras’ Wafer-Scale Engine architecture, which packs 44 GB of static random-access memory on each wafer-sized chip, the system eliminates the data movement bottlenecks common in GPU inference. This allows model weights to remain on-chip, enabling tokens to flow uninterrupted through model layers.

Benchmarks indicate the mode is significantly faster than competing models. According to data from Artificial Analysis, GPT-5.6 Sol on Ultrafast mode runs 11 times faster than Fable 5 and five times faster than Opus 4.8 on Fast mode. In tests using the 'Humanity's Last Exam' benchmark, which consists of 2,500 questions typically answerable only by PhDs, GPT-5.6 Sol Ultrafast completed the task in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5.

On the GDP-Val benchmark for economically valuable knowledge work, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation. Cerebras attributes these results to the architecture’s ability to scale smoothly with model size, paving the way for continued speed advantages on future frontier models. The company states that the technology is designed to accelerate time-sensitive, mission-critical work without forcing users to accept inferior results.

Potential use cases cited by Cerebras include root-causing production outages, responding to cyberattacks, and accelerating legal, financial, and engineering workflows. The firm suggests that faster inference allows organisations to put agents on the critical path of problems where every second counts, preserving customer trust and preventing lost revenue during high-stakes incidents.

This launch occurs against a backdrop of rapid growth in the AI sector, with both Google’s Gemini and OpenAI’s ChatGPT recently surpassing one billion monthly active users. The introduction of Ultrafast Mode signals a shift towards real-time processing capabilities for critical infrastructure and enterprise applications, aiming to provide a persistent edge for organisations using frontier AI.

Continue reading

More from Tech

Read next: Namecheap services disrupted by Phoenix data centre cooling failure
Read next: Trump administration authorises private firms to conduct cyber attacks on overseas criminals
Read next: Virgin Galactic opens public vote to name new Delta-class spaceship