Hasty Briefsbeta

Bilingual

AutoResearchExam: Measuring agents' ability to improve and generalize

21 days ago
  • AutoResearchExam is a benchmark of 29 open-ended ML research tasks measuring how quickly agents improve a private test score over 24 hours.
  • The evaluation separates visible progress (validation scores) from generalization (hidden-test AUARC) to detect overfitting.
  • Agents optimize validation scores and receive feedback but never see hidden test scores, with AUARC as the main metric.
  • Analysis shows models continue improving over 24 hours, with varying generalization gaps and hyperparameter tuning strategies.
  • Submission count alone does not correlate with higher scores; quality of research steps matters more.
  • Providing a research hint from prior trajectories boosts performance in shorter runs, especially for Claude Opus 5.
  • The Terminus 2 harness provides consistent evaluation across models, preserving rankings compared to native harnesses.