Atrocious AI-Written Tests
a day ago
- The author stopped reviewing tests because of the huge volume, but was forced to investigate after a comment change caused a test failure.
- That failing test was scanning source files with a regex for the keyword 'wall-clock', showing brittle and meaningless assertion behavior.
- Many tests run zero production code, such as checking an array of statuses does not contain 'failed' or 'cancelled' without testing any actual logic.
- Some tests mirror production code instead of testing it, so they only validate that the duplicate logic itself is consistent.
- Tests that assert static config values are low-value: they only fail when a config changes and don't protect against real regressions.
- Testing exact prompt headings is misguided because prompt effects on model output are nondeterministic; evals are the better mechanism.
- AI agents tend to add low-value tests to accompany code changes, even when the tests provide no meaningful safety.
- Of 6,429 total tests, 404 were low-value; they didn't slow the suite much but caused confusing failures and eroded trust in the test suite.