Entropic Thoughts
16 hours ago
- METR research shows LLMs pass tests more often than they produce mergeable code, with a stricter success criterion (maintainer approval) revealing much lower performance.
- The 50% success horizon dropped from 50 minutes to 8 minutes when using the more stringent criterion.
- Analysis of merge rates over time suggests no significant improvement since early 2025, despite a possible step-up in late 2024.
- Brier score comparisons indicate that a piecewise constant or constant function predicts merge rates better than a gentle upward slope, contradicting claims of steady improvement.
- The author questions the gap between perceived buzz and actual performance, noting that similar claims of progress in 2025 were not borne out by data.