- Medium implementation of a 2048 game with four browser files, ten engine tests, syntax checks, and a post-run evaluator, all dependency-free.
- Six controlled GPT-5.6-sol runs showed the same engineering loop skill lost on a small fix but succeeded on a medium build.
- Controlled experiment: same model, reasoning effort, starting commit, task contract, sandbox, and acceptance criteria; only repository-skill routing changed.
- Acceptance, required checks, and evidence completeness were prioritized; a cheaper failed run would not win.