Training loops in which a model (or another model) influences its own training during the forward pass, so capability gains compound without requiring proportionally more human-generated data or labels.
Tensions
If RSI lands, the implicit scaling law shifts from "more compute + more tokens" to "more compute + better feedback loops" — which favors labs with the largest coherent compute clusters and proprietary RL infrastructure, not whoever has the largest crawl. Whether RSI is the actual unlock or just a fundraising narrative is unsettled; Anthropic's pre-training hire of Karpathy is the most expensive bet on the bull case to date.