"As a Language Model": Chat Template Switches LLM Self-Referential Voice
6 hours ago
- LLMs often add disclaimers like "I'm just an AI" in self-referential responses, but the cause is unclear.
- The chat template acts as a switch: when present, it increases disclaimer voice and decreases experiential voice; when absent, the opposite occurs.
- A specific direction in model activations can steer this behavior; removing it reduces disclaimers, adding it increases them, while a random direction has little effect.
- Adding the disclaimer direction to instruct models without a chat template makes them disclaim as if the template were present.
- Researchers studying self-reports or introspection in models should control for the chat template as a confound, as model self-descriptions are not factual but partly determined by the template.