← Knowledge Base

Concept

Chinchilla Scaling

The DeepMind result that, for a fixed training compute budget, the optimal model size and training-token count grow at roughly equal rates (~20 tokens per parameter) — a law about training cost only.

Tensions

Once inference dominates lifetime cost (and RL is folded in), Chinchilla badly under-trains: frontier labs now appear to be pretraining ~100x past Chinchilla optimal because every saved inference FLOP pays off across hundreds of trillions of served tokens. Chinchilla is a starting point, not a target.

Related Concepts

cost equalization | scaling laws | inference economics

Last updated: May 17, 2026