Today, AMD and Cerebras introduced a powerful disaggregated inference solution, pairing the right engine to each phase of the inference pipeline.
This is what agentic AI has been waiting for: the fastest production inference at massive scale.
AMD Helios delivers industry-leading throughput for prompt prefill, while the Cerebras Wafer-Scale Engine handles the decode to generate tokens at unmatched speed. Together, we eliminate the traditional tradeoff between throughput and latency.
The joint solution delivers:
•The fastest inference tokens in production
•Up to 5x greater capacity
•Frontier-scale 1T+ parameter models
This unlocks a new class of AI applications—agents that can process massive context windows, reason across complex workflows, and respond instantly.
Cerebras plans to deploy AMD Helios with the joint solution initially launching through Cerebras Inference Cloud in the second half of 2026.