Ask a model if code is malicious and it reaches for its morals
4 hours ago
- Modern language models use a router to select a few experts per word; the selection depends on the question asked.
- When asked 'Is this code malicious?' the model routes through experts associated with morality more than vulnerability or legality.
- The model's workspace holds the concept of malware when asking about malice, not morality.
- Forcing the router to use another question's expert choices changes the model's answer.
- The morality path—experts weighted more by moral questions—is recruited by malice questions across models.
- Malice is the nearest neighbor to morality in routing distance on every model tested.
- Replacing the routing (transplant) moves the answer but does not transfer the donor's answer or concept.
- Pruning experiments show the morality path is what the router consults, not where judgment is stored.
- Conclusion: when asked about code maliciousness, models use moral machinery.