Hasty Briefsbeta

Bilingual

Is AI Progress Real? Four Independent Metrics Show It

a day ago
  • AI benchmark improvements often face issues like contamination, saturation, and reward hacking, making it hard to distinguish real progress from test-specific gains.
  • Four independent measures show a sharp acceleration in AI capabilities starting in late 2025, coinciding with releases like Grok 4.1, Gemini 3, Claude Opus 4.5, and GPT-5.2.
  • METR's time horizon metric grew from 4 minutes to 12 hours in under two years, indicating rapid advancement in task execution for frontier models.
  • TrackingAI's offline cognitive tests show AI systems went from outperforming fewer than 10% of humans to surpassing all but 1 in 44 within about two years.
  • Humanity's Last Exam scores jumped from 2.7% to 53% in a year, reflecting exponential compute increases needed for Ph.D.-level knowledge.
  • ARC-AGI-2 scores for abstract reasoning stalled for 15 months before surging in late 2025, with models like GPT-5.2 reaching 53% in weeks.
  • Four compounding mechanisms drive improvement: pretraining efficiency gains, reinforcement learning with verifiable rewards, harness engineering, and a flywheel effect from AI self-improvement.
  • Challenges like data exhaustion, reliability gaps in performance, and new benchmarks like ARC-AGI-3 (with interactive reasoning) could slow future progress.
  • The multiproxy test—similar to climate science's hockey stick—corroborates AI acceleration across diverse, independent metrics, suggesting real, accelerating improvement.