Can AI agents conduct open-ended AI research?
9 hours ago
- Shadow evaluations are proposed as a new method to measure AI agents' ability to conduct open-ended AI research.
- Agents managed to complete engineering tasks but failed to make substantial progress on research questions.
- Five recurring failure modes were identified: poor judgment, uncreative responses, ineffective backtracking, poor resource awareness, and instruction drift.
- A robustness check with a different model and scaffold reproduced the failures.
- The results suggest current AI agents can handle engineering but struggle with critical parts of the research lifecycle.