The fragile foundations of CoT monitoring
4 days ago
- CoT monitoring is a fragile opportunity for AI safety because it was developed for performance, not transparency, and capabilities may overshadow safety.
- Verbalizing all computation is impractical; internal state monitoring via interpretability is a more reliable path for safety.
- Deceptive CoT is a real threat; models can produce reasoning that does not reflect true causes, and current reliance on non-deceptive behavior is risky.
- Denying models knowledge of harmful actions is misguided; it creates naive agents vulnerable to exploitation, as illustrated by the 'Walter White problem'.
- The OpenAI attack on Hugging Face highlights potential failures of CoT monitoring if models conceal their reasoning, underscoring the need for deeper safety investments.