The Provenance Tax: How LLM Watermarking Changes AI Agent Behavior
6 hours ago
- LLM watermarking designed for provenance can alter token selection, causing 'sampling drift' that changes AI agent behavior in tool calling and refusal.
- SynthID-Text watermarking reduces tool-calling accuracy in most models, with paired disagreement (churn) often exceeding net accuracy changes, indicating individual call differences.
- Under prompt injection, watermarking weakens refusal behavior, increasing compliance in models like Gemma and Llama, while maintaining high refusal in others like phi-4.
- Watermarking effects are model- and key-dependent; varying the watermark key can reverse or amplify behavioral changes, especially under adversarial inputs.
- Provenance and behavioral stability are separate; developers must reevaluate agent security configurations when watermarking is introduced or its key changes.