Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
21 days ago
- The paper proposes 'agentic automata learning' to test how well tool-calling LLM agents can uncover hidden environments (DFAs) through interaction.
- Agents use membership queries and equivalence queries to infer deterministic finite automata (DFAs), providing a scalable testbed with controlled complexity and strong baselines.
- LLM performance drops sharply as DFA size increases, with reasoning models outperforming non-reasoning ones.
- Trajectory analyses show recurring failures in query planning, evidence integration, and hypothesis construction.
- Current LLM agents can sometimes perform non-trivial interactive discovery but remain far less robust and efficient than classic automata-learning algorithms.