Cerebras Doubles Performance With WSE-3 Turbo and Launches First Rack-Scale CS-4 System
Cerebras has announced a major upgrade to its artificial intelligence hardware lineup, introducing the WSE-3 Turbo processor and its first true rack-scale system, the CS-4. The WSE-3 Turbo doubles the performance of the original WSE-3 by running at twice the clock speed, delivering 250 petaFLOPS (PFLOPS) of computing power for sparse 16-bit floating-point data. The new CS-4 system houses three of these processors in a single rack, enabling six times the performance of a single WSE-3 system and positioning Cerebras to compete directly with graphics processing unit (GPU) makers and other AI chip companies.
What Makes the WSE-3 Turbo Different From Previous Versions?
Cerebras has been building wafer-scale engines since 2019, using an unconventional approach: instead of slicing a silicon wafer into individual chips, the company uses almost the entire wafer as a single processor. This allows for a much larger chip with more computing cores packed onto it. The WSE-3 Turbo continues this tradition but with a significant speed boost.
The Turbo version is not an entirely new design. Rather, it takes the existing WSE-3 architecture and cranks up the clock speeds across every component. The processor still contains 900,000 AI cores and 44 gigabytes of on-processor memory, and it is still manufactured on TSMC's 5-nanometer process node. However, nearly every critical performance metric has doubled:
- Computing throughput: Increased from 125 PFLOPS to 250 PFLOPS for sparse 16-bit floating-point operations
- Memory bandwidth: Doubled from 21.6 petabytes per second to 43.2 petabytes per second, allowing faster data movement within the chip
- Fabric bandwidth: Increased from 26.8 petabytes per second to 53.5 petabytes per second for internal communication between cores
- Network bandwidth: Doubled from 150 gigabytes per second to 300 gigabytes per second for connecting to external systems
The power consumption likely doubled as well, bringing it to around 54 kilowatts per processor. Cerebras achieved this by carefully managing the voltage and frequency curve, avoiding the superlinear power increases that typically occur when pushing chips to higher speeds.
How Does the CS-4 Rack System Scale AI Computing?
The CS-4 represents Cerebras's first true rack-scale system, a significant departure from its previous approach. The earlier CS-3 system housed a single WSE-3 processor in a 16-unit liquid-cooled server, and even when multiple CS-3 systems were deployed together, they did not function as an integrated, scale-up system. The CS-4 changes this fundamentally.
A full CS-4 rack now contains three WSE-3 Turbo processors working together as a unified system. This represents a 50 percent increase in wafer-scale engines compared to the previous CS-3 Rack configuration. Combined with the Turbo's doubled performance, a single CS-4 rack delivers approximately six times the performance of a single CS-3 system.
To house these faster, more power-hungry processors, Cerebras redesigned the entire rack architecture. The company introduced what it calls the Nexus platform, a modular design that places power supplies, cooling fans, and supporting hardware at the front of the rack while positioning the wafer-scale engines at the rear. This modular approach is intended to support multiple generations of processors and technologies, allowing Cerebras to upgrade components without completely redesigning the system.
How to Understand Cerebras's Competitive Position in AI Infrastructure
- Scale-up vs. scale-out: Cerebras is positioning the CS-4 as a scale-up system, where multiple processors work together tightly within a single rack. This contrasts with the scale-out approach used by GPU makers, where many independent systems are networked together. Scale-up systems can offer lower latency and tighter integration for certain AI workloads.
- Wafer-scale advantage: By using nearly an entire silicon wafer as a single processor, Cerebras maximizes the number of computing cores and on-chip memory available. This reduces the need to move data off-chip, which is often a bottleneck in AI computing. The 44 gigabytes of on-processor memory is substantially larger than what individual GPU cores can access locally.
- Rack-scale architecture: The CS-4's ability to integrate three processors per rack with optimized networking and power delivery allows Cerebras to compete with recent announcements from Nvidia and AMD in the rack-scale AI infrastructure space. This is a direct response to the industry's shift toward larger, more integrated systems.
Cerebras's timing aligns with a broader industry trend. As AI models grow larger and more complex, companies are moving away from loosely coupled clusters of individual accelerators toward tightly integrated systems that can handle massive workloads more efficiently. The CS-4 and WSE-3 Turbo represent Cerebras's answer to this shift, offering a fundamentally different architecture from the GPU-based systems that currently dominate the market.
The company has been steadily improving its wafer-scale engine technology since the original 2019 launch. Each generation has brought higher performance and attracted more developer interest and sales momentum. The WSE-3, launched in 2024, already offered 125 PFLOPS of sparse 16-bit performance. Now, with the Turbo version doubling that performance and the CS-4 enabling three processors per rack, Cerebras is making a clear statement about its ambitions in the AI infrastructure market.
The broader context matters here. The AI market continues to boom, and companies are racing to build systems that can train and run increasingly large language models and other AI applications. Cerebras is not trying to replace GPUs entirely but rather to position its wafer-scale engines as a more capable alternative for certain workloads, particularly those that benefit from large amounts of on-chip memory and high internal bandwidth. The CS-4 and WSE-3 Turbo represent the company's most ambitious effort yet to compete at scale.