Cerebras CS-4 Supports 10T Models and Doubles Token Rates
+Cerebras has updated its AI-compute rack. Due to ship later this quarter, the CS-4 rack roughly doubles the throughput of the earlier CS-3 to 4,465 tokens/sec. Performance gains come from an updated “chip,” the Wafer-Scale Engine 3 Turbo and lower-latency communication among wafers. More importantly, the latter helps the Cerebras design to support large (e.g., 10 trillion-parameter) models while sustaining performance.
The rack architecture has radically changed. Now oriented vertically, wafers attach to the rack using backpacks, small chassis attached to the back of the rack. Power and I/O subsystems are also modular and slide into the front of the system. Scale-out I/O is via RoCE, helping the CS-4 fit into heterogeneous designs that pair it with another system type. For example, Amazon and AMD have announced deployments combining Trainium and Helios/Instinct, respectively, with Cerebras boxes and splitting LLM prefill and decode between the types.
The CS-3 distinguishes itself from alternative AI accelerators through its leading per-use token rates. The CS-4 extends this advantage and also raises overall tokens-per-second throughput, where GPUs have held a vast lead. The new design, therefore, positions Cerebras to expand beyond its low-latency premium-user niche.
The Cerebras approach has had three main disadvantages:
1. Software compatibility. Theoretically, the CS/WSE can support any model, but in practice, the company has hosted a selection of models in its own clouds or has sold to customers running a limited number (e.g., 1) model.
2. Memory capacity. Again, there’s no unusual limit theoretically. In practice, Cerebras hasn’t run the biggest models or supported the largest context sizes.
3. Overall weirdness of a wafer-scale chip, necessitating customers buy a whole system instead of integrating the WSE into racks conforming with their other infrastructure.
Issue #1 remains. The new wafer-wafer I/O and “Turbo” WSE mitigates Issue #2. As for Issue 3, AI-accelerator suppliers have increasingly forward-integrated, supplying complete rack blueprints, if not the actual racks. That is, the market has moved in the direction of Cerebras. The modular design of the CS-4 mitigates the weirdness factor by improving maintainability. Nonetheless, Cerebras remains an odd duck. But weirdness is tolerable when it delivers a premium user experience in the form of blisteringly fast token rates.
Other contents