The initial large-scale self-supervised training of a base model on broad data, setting raw capability before any post-training.
Tensions
Two 2026 views collide. Yao Shunyu argues pre-training has not plateaued and that most people who think a scaling law is exhausted actually have an undiscovered bug. Luo Fuli argues the pre-training gap across frontier labs has essentially closed, so last era's pre-training success no longer guarantees this era's lead, and a roughly 1T base model is now just the entry ticket.