- CrucibleBench evaluates language models in a persistent MUD text world over 50 turns with hidden social objectives.
- It uses 'lateral thinking with withered technology'—applying mature, inexpensive tech (MUDs) to measure AI behavior.
- Key constraints include an enumerable action space, explicit social feedback from NPCs with trust/suspicion states, and within-run persistence.
- The central finding: an LLM judge component can reorder the leaderboard by up to six positions while aggregate reliability statistics stay silent.