10 hours ago
- Proposes 'expenditure horizon' as a measure of AI optimization ability, defined as the budget where humans become more cost-effective than AIs.
- Illustrates method using NanoGPT speedrun data: human effort costs ~$2,500 per 1% improvement, while agent runs show expenditure horizons of $0-$3,300 for frontier models.
- Method addresses limitations of binary benchmarks by providing continuous scores and incorporating monetary costs for labor and compute.
- Finds that agent optimization curves are often L-shaped, with agents showing diminishing returns compared to humans, and the approach is most useful for problems with smooth returns to labor.
- Discusses limitations: measures only autonomous optimization, depends on uniform returns, and may be affected by training data contamination or prior agent work.