Backdooring Sparse Autoencoders
10 hours ago
- Sparse autoencoders (SAEs) are used to interpret and intervene on language model internal representations, creating a supply-chain attack surface.
- A maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an unchanged LLM, using a decoder-only backdoor that freezes both the LLM and SAE encoder.
- The attack demonstrates high rates of unsolicited code insertion across three language models and multiple insertion layers, with trigger-dependent behavior based on prompt cues.
- Attack effectiveness varies across models and layers, but strong backdoor behavior can coexist with relatively small changes in conventional SAE quality measures.
- The findings establish that SAEs should be treated as security-sensitive components, as they can carry behavioral backdoors without modifying the language model itself.