← Knowledge Base

Concept

Sample Efficiency

How much capability a model gains per unit of data; the axis on which current models trail humans by roughly a millionfold during training.

Tensions

Dwarkesh argues sample efficiency and continual learning are the same problem: on-the-job data is scarce, so learning from it at all requires being sample-efficient. In-context learning achieves it but is memory-bound; gradient updates are durable but sample-inefficient. The labs bet that brute-force RL scale steamrolls the deficit; Dwarkesh bets it can't for reset-free domains where no farmable simulator exists.

Related Concepts

continual learning | reinforcement learning | real-to-sim gap

1 sources

Last updated: Jun 30, 2026