Cerebras Unveils CS-4 Wafer-Scale System Claiming 30x Inference Speed
The CS-4 utilises a wafer-scale chip and integrated power delivery to promise significantly higher throughput per watt compared to conventional GPU systems.
Cerebras has officially launched the CS-4, a rack-scale artificial intelligence system built around a wafer-scale chip architecture. The company positions the new hardware as the first iteration of its Cerebras Nexus Platform, designed to replace conventional GPU clusters with a single-chip solution that claims to deliver inference speeds up to 30 times faster than existing production systems.
Powered by the WSE-Turbo chip, the CS-4 architecture is engineered to improve efficiency alongside raw speed. According to manufacturer data, the system delivers up to 10 times more throughput per watt than the previous CS-3 generation. This performance gain is attributed to a new modular design that integrates compute, power, and input/output capabilities into a unified structure, aiming to simplify manufacturing and maintenance for hyperscale datacentre operators.
A key innovation in the CS-4 is the 'Wafer-Scale Backpack', a compact 3D package that folds the wafer, power conversion, direct liquid cooling, and high-speed I/O into a single assembly. This design reduces the component count by 50 per cent and positions power delivery just 0.5 millimetres from the processor. This proximity is roughly 100 times closer than the 50-millimetre distance found on conventional GPU boards, a move Cerebras states nearly eliminates board-level power loss and allows for higher operating frequencies.
The system also introduces a new programmable I/O subsystem that doubles bandwidth and reduces latency. By enabling wafers to be linked within and across racks without a switch, the CS-4 achieves wafer-to-wafer interconnect latency of 2 microseconds. This low-latency connection allows the system to process more than 1,000 tokens per second on models exceeding 10 trillion parameters, maintaining interactive decode performance at unprecedented scales.
To further accelerate deployment, Cerebras has separated the stable power, cooling, and network layer from the compute hardware. The Cerebras PowerRack can be installed and qualified by facility teams before the compute hardware arrives. Once the PowerRack is ready, the modular compute backpacks slide into place and connect to the existing infrastructure, a process the company says reduces deployment time from days to hours while simplifying future upgrades.
Pricing and specific availability details for the CS-4 were not included in the source material. Performance metrics cited, including the 30x speed increase and 10x throughput improvement, are based on manufacturer claims without independent verification in the provided text.

