Cerebras launches CS-4 rack system with WSE-3 Turbo, claims 30x GPU inference speed
Cerebras on Aug 18 launched CS-4 rack with 3 WSE-3 Turbo engines at 750 PFLOPS and 129.6 PB/s memory bandwidth; in GPT-OSS-120B testing each user generates 4,465 tokens per second, marketed as 30x GPU; shipping in Q3.
Wafer-scale AI inference accelerator maker Cerebras unveiled its WSE-3 Turbo chip and CS-4 rack system on August 18, 2026. WSE-3 Turbo integrates 900,000 cores, 44GB SRAM cache and 4 trillion transistors, with FP16 sparse AI compute, memory bandwidth, on-chip interconnect and I/O bandwidth all doubled versus the prior WSE-3.The CS-4 rack integrates three WSE-3 Turbo chips, delivering 750 PFLOPS of AI compute and 129.6 PB/s of memory bandwidth. In GPT-OSS-120B testing, CS-4 generates 4,465 tokens per second per user, marketed as 30x the speed of GPU solutions and a 93% improvement over CS-3. CS-4 supports models of more than 50 trillion parameters.CS-4 uses the new Nexus rack platform, with compute, power and interconnect subsystems redesigned. Component count is reduced by 50% and manufacturing automation is substantially increased, with near-elimination of board-level power loss to deliver double the power to WSE-3 Turbo.CS-4 natively supports inference load splitting, allowing deployment as a decode unit alongside heterogeneous prefill hardware. Shipment is expected to begin in Q3 2026.