Learning in non-stationary environments that cannot be reset and replayed in parallel — where each rollout interacts irreversibly with the real world and the reward may take months or years to materialize.
Tensions
A known-open problem in RL. Most economically valuable skills (running a business, litigation, politics, trading) are reset-free, so the farmable-simulator trick that works for coding — and eventually computer use — doesn't apply. That forces reliance on sample efficiency rather than parallel grinding.