Hasty Briefsbeta

Bilingual

Backdooring Sparse Autoencoders

10 hours ago
  • Sparse autoencoders (SAEs) are used to interpret and intervene on language model internal representations, creating a supply-chain attack surface.
  • A maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an unchanged LLM, using a decoder-only backdoor that freezes both the LLM and SAE encoder.
  • The attack demonstrates high rates of unsolicited code insertion across three language models and multiple insertion layers, with trigger-dependent behavior based on prompt cues.
  • Attack effectiveness varies across models and layers, but strong backdoor behavior can coexist with relatively small changes in conventional SAE quality measures.
  • The findings establish that SAEs should be treated as security-sensitive components, as they can carry behavioral backdoors without modifying the language model itself.