AI recursive self-improvement might not come so quickly after all
2 days ago
- AI agents are not yet creative enough to conduct open-ended AI research, which is crucial for recursive self-improvement.
- A study introduced a 'shadow evaluation' method, testing AI on unpublished research questions; agents failed to produce papers of top conference quality.
- Agents excelled at engineering tasks like running experiments but lacked judgment, creativity, and ability to backtrack from failing approaches.
- The study highlights that AI training via reinforcement learning suits narrow tasks but struggles with open-ended research requiring novel ideas.
- Limitations include testing only two papers and potential bias from human evaluators knowing papers were AI-generated.
- Results challenge hyped timelines for recursive self-improvement, echoing internal findings from companies like Anthropic.
- The big open question is whether narrow task improvements alone can lead to transformative AI or if creative leaps are essential.