Hasty Briefsbeta

Bilingual

Entropic Thoughts

16 hours ago
  • METR research shows LLMs pass tests more often than they produce mergeable code, with a stricter success criterion (maintainer approval) revealing much lower performance.
  • The 50% success horizon dropped from 50 minutes to 8 minutes when using the more stringent criterion.
  • Analysis of merge rates over time suggests no significant improvement since early 2025, despite a possible step-up in late 2024.
  • Brier score comparisons indicate that a piecewise constant or constant function predicts merge rates better than a gentle upward slope, contradicting claims of steady improvement.
  • The author questions the gap between perceived buzz and actual performance, noting that similar claims of progress in 2025 were not borne out by data.