Splitting an inference request across heterogeneous chips — typically Nvidia GPUs at the center for compute-bound work, with specialized accelerators (Groq, Cerebras) for memory-bandwidth-bound decode in front of them.
Tensions
Disaggregation extends the useful life of older GPUs (10-15 years as decode-front nodes drop the workload they're bad at), which collides with the bear thesis that GPU lifespans are ~2 years. The amortization implication — cheaper financing on the most depreciation-resistant component — is itself a competitive moat that pure-play ASIC clouds don't get.
Related Concepts
inference economics | hardware | batching | memory bandwidth