Hasty Briefsbeta

Bilingual

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

6 hours ago
  • LLMs often add disclaimers like "I'm just an AI" in self-referential responses, but the cause is unclear.
  • The chat template acts as a switch: when present, it increases disclaimer voice and decreases experiential voice; when absent, the opposite occurs.
  • A specific direction in model activations can steer this behavior; removing it reduces disclaimers, adding it increases them, while a random direction has little effect.
  • Adding the disclaimer direction to instruct models without a chat template makes them disclaim as if the template were present.
  • Researchers studying self-reports or introspection in models should control for the chat template as a confound, as model self-descriptions are not factual but partly determined by the template.