- MirrorCode is a benchmark co-developed with METR for long-horizon AI coding tasks, requiring complete end-to-end reimplementation of programs without access to the original source code.
- The benchmark features 25 target programs spanning Unix utilities, bioinformatics, cryptography, interpreters, compression, and other domains.
- It provides scale-aware evaluations with realistic inference budgets (up to $2,600 and 19 days per run), and is sandboxed and cheat-resistant with held-out tests.
- Frontier models can solve complex tasks; for example, Claude Opus 4.7 reimplemented the gotree bioinformatics toolkit in 14 hours at a cost of $251, a task estimated to take a human engineer 2–17 weeks.