- AMD and Cerebras Systems announced a technical partnership to deliver a new disaggregated AI inference solution.
- The solution combines AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine for ultra-low-latency and high throughput.
- The joint solution is expected to deliver up to 5x higher tokens per second per watt (T/s/W) compared to a Cerebras WSE-only configuration.
- AMD Helios provides high-throughput prompt processing, while Cerebras WSE accelerates memory-bandwidth-intensive token generation with ultra-low latency.
- The solution addresses the growing diversity of AI inference workloads with heterogeneous infrastructure.
- Cerebras plans to deploy AMD Helios in its data centers, with availability through Cerebras Cloud in the second half of 2026.