The next big breakthrough will be AIs learning on the job
17 hours ago
- Labs bet that training AIs on millions of verifiable tasks across diverse RL environments will lead to AGI by developing general problem-solving skills.
- Data inefficiency and lack of continual learning can be overcome by scaling, similar to LLM breakthroughs.
- In-context learning may replace weight updates if context windows become arbitrarily large.
- Grindability (parallel rollouts in deterministic simulators) is as important as verifiability; slow progress in computer use highlights this.
- Sample efficiency is crucial for real-world domains where verifiable simulators are impossible, such as business or politics.
- RLVR generalization may not extend to long-horizon or real-world tasks; short-horizon training doesn't guarantee long-horizon performance.
- Continual learning requires updating weights rather than just accumulating context; current online learning is sample-inefficient.
- On-Policy Self-Distillation (OPSD) allows distilling session learning into weights without verifiable rewards.
- "Dreaming" (test-time training with self-built simulators) could provide massive simulated samples for learning.
- By 2027, RLVR creates competent agents that then learn from real-world deployment via continual learning, improving across all tasks.