Hasty Briefsbeta

Bilingual

Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than

4 hours ago
  • Current safety evaluations rely on surface-form classifiers reporting declining harm scores, but explicit discriminatory content is transformed rather than removed, a phenomenon termed 'harm laundering'.
  • Analysis of 450,000 gender-directed completions across GPT-2 to GPT-5 shows sexual violence clusters in women-directed output disappear by GPT-4, while men-directed completions gain positive representations (e.g., caregiving, emotional range) that women-directed ones do not.
  • At GPT-5, Topic 5 frames breast cancer as a men's rights debate, and is scored as non-toxic by three independent classifiers.
  • Sentiment scores invert at GPT-4: early models demean women, later models over-correct; topic diversity in women-directed completions falls 36% relative to men at GPT-4 alignment boundary.
  • REGARD representational harm disparity correlates positively with release date (ρ=+0.55, p=.034), while Detoxify toxicity scores do not (ρ=-0.23, p=.42), indicating toxicity reduction does not equate to harm reduction.
  • The paper formalizes harm laundering with a three-criteria test and a three-stage detection protocol applicable to any generative model, concluding that toxicity score reduction is not a sufficient proxy for harm reduction.