Everything done to a base model after pre-training (SFT, RLHF, RL on agent trajectories) to shape its behavior and unlock capability.
Tensions
As compute shifts toward post-training (Luo Fuli reports top teams moving from a roughly 3:5:1 to a 3:1:1 research/pre-train/post-train split, with pre-train to post-train heading to 1:1), the lead increasingly comes from RL and agent infra rather than raw pre-training scale. This also reshapes org design: post-training now rewards background diversity, and moving pre-training people into post-training is a strong complement.