Cerebras Systems said on October 1, 2026 that it increased inference throughput by 5x in early results using a technique called disaggregation, with the same number of Cerebras systems and no loss in token generation speeds. The disclosure came in Disaggregated Inference From the Ground Up, a company blog post by Isaac Tai and Zhenwei Gao that opens a planned series on the subject. The post frames the series for readers who have heard the term disaggregation, or the claim that prefill is…