AMD and Cerebras partner on low-latency, high-throughput AI inference — EPYC processors in Helios rack-scale infrastructure paired with Cerebras’ Wafer-Scale Engine (WSE) solutions

AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD’s EPYC processors in Helios rack-scale infrastructure with Cerebras’ Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD’s CPUs and Instinct GPUs with high throughput of Cerebras’s Wafer Scale Engines (WSE) processors.

AMD and Cerebras expect the new inter-rack-scale platform — based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside — to be responsible for prompt processing and large context windows, whereas Cerebras’ WSE will take care of the memory-bandwidth-intensive token-generation stage.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *