The author discovered that lazy evaluation of research papers works: by writing without citations, readers will send references.
Bayesian spam filtering was known since 1998, but early filters (Pantel & Lin) had poor performance (92% catch, 1.16% false positives), while the author's later filter achieved 99.5% catch with <0.03% false positives.
Five key differences explain the performance gap: training on too little data, ignoring message headers, stemming tokens, using all tokens instead of the 15 most significant, and not biasing against false positives.
Tokenization improvements include preserving case, treating exclamation points as constituents, marking tokens from headers/URLs, and using degeneration to handle unknown tokens by falling back to less specific versions.
A larger vocabulary (187,000 tokens vs. 23,000) increases discrimination but risks misses; degeneration helps by looking up progressively less specific tokens.
To handle equally interesting tokens, the author scales probabilities for tokens that occur only in one corpus (e.g., .9999 for >10 occurrences, .9998 otherwise).
HTML handling should be moderate: ignore most tags but parse links, images, and font tags; future spams may avoid HTML entirely.
Over one month, the filter caught 99.75% of spam (4 missed out of 1750) but had 3 false positives out of ~7740 legitimate emails (0.06% rate).
False positives are treated as bugs, not performance metrics; they often come from newsletters, company emails, or unusual formatting (e.g., all uppercase).
Future work focuses on two hard-to-filter spam types: personal ad spams (using neutral language) and outsourcing spams from companies (e.g., Bulgaria) that use programming jargon.
Individual user-specific filters are advantageous because they create diverse probability profiles, slowing spammers' ability to optimize their attacks.
Network-level filters are less promising because spammers can bypass them, but individual filters can adapt to each user's mail and make spammers' testing cycles inefficient.