- Benchmarks are often perceived as authoritative and objective but may lack depth if not scrutinized.
- The author critiques AI Code Review Bench methodology, arguing it lacks clear problem definition and splits AI code review into two problems: human assistance and machine verification.
- Human-focused code review prioritizes recommendations to optimize limited attention, while machine-focused code review emphasizes exhaustive analysis for repair agents.
- The paper acknowledges challenges like reviewer imperfections and Goodhart's Law but may overemphasize proxies like agreement with human reviewers over outcomes like reducing production failures.