Is AI Progress Real? Four Independent Metrics Show It
a day ago
- AI benchmark improvements often face issues like contamination, saturation, and reward hacking, making it hard to distinguish real progress from test-specific gains.
- Four independent measures show a sharp acceleration in AI capabilities starting in late 2025, coinciding with releases like Grok 4.1, Gemini 3, Claude Opus 4.5, and GPT-5.2.
- METR's time horizon metric grew from 4 minutes to 12 hours in under two years, indicating rapid advancement in task execution for frontier models.
- TrackingAI's offline cognitive tests show AI systems went from outperforming fewer than 10% of humans to surpassing all but 1 in 44 within about two years.
- Humanity's Last Exam scores jumped from 2.7% to 53% in a year, reflecting exponential compute increases needed for Ph.D.-level knowledge.
- ARC-AGI-2 scores for abstract reasoning stalled for 15 months before surging in late 2025, with models like GPT-5.2 reaching 53% in weeks.
- Four compounding mechanisms drive improvement: pretraining efficiency gains, reinforcement learning with verifiable rewards, harness engineering, and a flywheel effect from AI self-improvement.
- Challenges like data exhaustion, reliability gaps in performance, and new benchmarks like ARC-AGI-3 (with interactive reasoning) could slow future progress.
- The multiproxy test—similar to climate science's hockey stick—corroborates AI acceleration across diverse, independent metrics, suggesting real, accelerating improvement.