Julie Choi, Cerebras
Julie Choi of Cerebras Systems joins John Furrier of theCUBE Research to discuss a collaboration between Cerebras and AMD that advances disaggregated artificial intelligence inference and ultra-low-latency decode performance. Choi outlines the Cerebras wafer-scale engine and the integration with AMD Helios. They explain how a disaggregated inference architecture pairs Helios for compute-intensive pre-fill with Cerebras for memory-bandwidth constrained low-latency decode and they examine the technical rationale and implications for agentic coding and multimodal workloads. Choi states the combined Helios–Cerebras solution delivers up to 5x higher tokens per second per watt. They describe Helios managing compute-intensive pre-fill while Cerebras addresses the memory-bandwidth constrained low-latency decode. Choi adds that the companies co-develop engineering and go-to-market plans, with Helios deployments planned in Cerebras data centers and broader rollout targeted before year-end. The discussion highlights implications for AI inference efficiency, agentic systems, multimodal models and developer coding workflows. This segment provides practical insights for technology leaders and infrastructure teams evaluating disaggregated inference, low-latency decoding and energy-efficient AI deployment.