← Knowledge Base

Concept

Post-training

Everything done to a base model after pre-training (SFT, RLHF, RL on agent trajectories) to shape its behavior and unlock capability.

Tensions

As compute shifts toward post-training (Luo Fuli reports top teams moving from a roughly 3:5:1 to a 3:1:1 research/pre-train/post-train split, with pre-train to post-train heading to 1:1), the lead increasingly comes from RL and agent infra rather than raw pre-training scale. This also reshapes org design: post-training now rewards background diversity, and moving pre-training people into post-training is a strong complement.

Related Concepts

pre-training | reinforcement learning | agent

Last updated: May 29, 2026