"Uncensored" open LLMs are measurably more optimistic than their base models
a day ago
- Abliteration of refusal direction in model weights leads to measurable side effects beyond refusal removal, as shown by systematic behavioral shifts in a decision-making task that elicits no refusals.
- Abliterated models from two Mixture-of-Experts families (Gemma-4-26B-A4B-it and Qwen3-30B-A3B-instruct) showed increased optimism (higher positivity in decisions), longer justifications, and fewer uncertainty words in self-critiques.
- The effect on expressed confidence was opposite across families: Gemma models became less confident, while Qwen models became more confident.
- The ablation process did not degrade instruction-following ability, and neither abliterated nor base models demonstrated economic skill in stock predictions.
- Two contamination channels were detected in community-modified checkpoints, suggesting toolchain artifacts are common in such studies.
- Deploying an 'uncensored' model as an agent means deploying a measurably different decision-maker, not the base model without refusals.