Curated, verifiable task environments in which humans build the framework, tools, and reward rubrics so models can learn to perform real, multi-tool, long-horizon work; the dominant new post-training data type in 2026.
Tensions
Foody argues RL environments could subsume the entire economy because the addressable market equals what humans can do that models cannot, and each new layer of tool or trajectory complexity reopens that gap. Skeptics counter with near-term plateau fears and the question of how long humans stay in the loop before superintelligence closes it.
Related Concepts
reinforcement learning | evaluation | post-training | real-to-sim gap