AutoResearchExam: Measuring agents' ability to improve and generalize
21 days ago
- AutoResearchExam is a benchmark of 29 open-ended ML research tasks measuring how quickly agents improve a private test score over 24 hours.
- The evaluation separates visible progress (validation scores) from generalization (hidden-test AUARC) to detect overfitting.
- Agents optimize validation scores and receive feedback but never see hidden test scores, with AUARC as the main metric.
- Analysis shows models continue improving over 24 hours, with varying generalization gaps and hyperparameter tuning strategies.
- Submission count alone does not correlate with higher scores; quality of research steps matters more.
- Providing a research hint from prior trajectories boosts performance in shorter runs, especially for Claude Opus 5.
- The Terminus 2 harness provides consistent evaluation across models, preserving rankings compared to native harnesses.