← Knowledge Base

Concept

Latency

The wall-clock time to produce a token, with a hard lower bound set by total parameters divided by aggregate HBM bandwidth in the scale-up domain.

Tensions

Latency floors and cost floors live on opposite ends of the batch-size curve. Pipelining adds a few milliseconds per rack hop and stacks across pipeline stages, so pushing scale-up larger (rather than scale-out) is the only path to lower latency at frontier model sizes.

Related Concepts

memory bandwidth | batching | scale-up domain | pipelining | inference economics

Last updated: May 17, 2026