a month ago
- Critiques adaptations of the mirror test for LLMs as flawed because they translate visual tests into text.
- Proposes a better analogy: modify an LLM's own textual output subtly and see if it notices the anomaly.
- Describes an experiment with Gemma 4 31B-IT where corrupted text (replacing 'g' with 'sg') was introduced.
- Gemma spontaneously detected the corruption in its thinking trace, shifting from first-person to third-person language.