← Knowledge Base

Concept

Pre-training

The initial large-scale self-supervised training of a base model on broad data, setting raw capability before any post-training.

Tensions

Two 2026 views collide. Yao Shunyu argues pre-training has not plateaued and that most people who think a scaling law is exhausted actually have an undiscovered bug. Luo Fuli argues the pre-training gap across frontier labs has essentially closed, so last era's pre-training success no longer guarantees this era's lead, and a roughly 1T base model is now just the entry ticket.

Related Concepts

post-training | scaling laws | Chinchilla scaling

Last updated: May 29, 2026