Hasty Briefsbeta

Bilingual

METR introduces Expenditure Horizon metric

8 hours ago
  • Proposes 'expenditure horizon' as a measure of AI optimization ability, defined as the budget where humans become more cost-effective than AIs.
  • Illustrates method using NanoGPT speedrun data: human effort costs ~$2,500 per 1% improvement, while agent runs show expenditure horizons of $0-$3,300 for frontier models.
  • Method addresses limitations of binary benchmarks by providing continuous scores and incorporating monetary costs for labor and compute.
  • Finds that agent optimization curves are often L-shaped, with agents showing diminishing returns compared to humans, and the approach is most useful for problems with smooth returns to labor.
  • Discusses limitations: measures only autonomous optimization, depends on uniform returns, and may be affected by training data contamination or prior agent work.