Hasty Briefsbeta

Bilingual

Atrocious AI-Written Tests

a day ago
  • The author stopped reviewing tests because of the huge volume, but was forced to investigate after a comment change caused a test failure.
  • That failing test was scanning source files with a regex for the keyword 'wall-clock', showing brittle and meaningless assertion behavior.
  • Many tests run zero production code, such as checking an array of statuses does not contain 'failed' or 'cancelled' without testing any actual logic.
  • Some tests mirror production code instead of testing it, so they only validate that the duplicate logic itself is consistent.
  • Tests that assert static config values are low-value: they only fail when a config changes and don't protect against real regressions.
  • Testing exact prompt headings is misguided because prompt effects on model output are nondeterministic; evals are the better mechanism.
  • AI agents tend to add low-value tests to accompany code changes, even when the tests provide no meaningful safety.
  • Of 6,429 total tests, 404 were low-value; they didn't slow the suite much but caused confusing failures and eroded trust in the test suite.