Back to news
launchCerebrasOpenAIAMDAWS2026-08-19

Cerebras unveils CS-4 rack system, claims 30x GPU inference speed

On August 18 Cerebras launched CS-4 at Supernova 2026 with three WSE-3 Turbo wafers per rack, announcing partnerships with OpenAI, AMD, AWS, Arista and more.

On August 18, AI chip company Cerebras Systems launched its next-generation rack-scale AI system CS-4 at SUPERNOVA 2026 in San Francisco, claiming it is "the industry's fastest AI accelerator." CS-4 is the first product of Cerebras' new Nexus rack-scale platform, driven by three Wafer Scale Engine 3 Turbo (WSE-3T) wafer-scale processors.

WSE-3T is the largest AI processor ever built, with each chip manufactured on TSMC's 5nm process at 46225 mm², integrating 4 trillion transistors and 900,000 AI-optimized cores with 44GB SRAM. The CS-4 system delivers 750 PFLOPS of AI compute, 129.6 PB/s memory bandwidth, and 7.2 Tb/s I/O throughput, twice the speed of the prior CS-3. In GPT-OSS-120B comparison, CS-4 generates over 4400 tokens per user per second, up to 30x GPU systems, supporting models with more than 50 trillion parameters.

Cerebras simultaneously announced partnerships with OpenAI, AMD, Arista Networks, AWS, Figma, Cognition, and CrowdStrike. The AMD and AWS Trainium collaborations focus on disaggregated inference, with Helios or Trainium handling prefill and CS-4 handling ultra-fast token decoding, achieving 5x gain over pure WSE configurations and up to 10x over GPU systems.

The physical Nexus platform redesign brings significant efficiency gains, with a front-mounted power section and rear-mounted "Wafer-Scale Backpack" modules that cut component count by 50% versus CS-3 and shrink deployment time from days to hours. CS-4 shipments begin in Q3 2026, with the company targeting roughly 2x annual performance improvement and 20x throughput by end of 2027.

CerebrasCS-4WSE-3T推理机架级