Why models write slop: the environments are too small
2 days ago
- Data scaling is far behind compute scaling because generating high-quality data remains a difficult, bespoke task.
- Anthropic's edge comes from focusing on coding models, which allows it to collect vast amounts of labeled data from users who pay for the privilege.
- Programming is a key domain, but achieving recursive self-improvement (RSI) requires more than just code data.
- Labs are trying to scale data generation, but it remains a cottage industry compared to industrial-scale compute.
- Models produce slop because their training environments are too small and fail to teach critical capabilities like taste and long-term planning.
- Games (e.g., Factorio, Civilization) provide rich environments that could let models develop missing skills through trial and error, potentially unlocking RSI.
- The ultimate goal is to create self-sustaining AI agents that can collaborate on solving big challenges.