- HANDBOOK.md is a benchmark of 65 agentic tasks based on enterprise employees following company handbooks.
- Each task places an agent in a self-contained company environment with a long standard operating procedure (20–124 pages).
- Tasks span five domains: finance, medical billing, insurance, logistics, and HR, across ten fictional companies.
- Grading is fully deterministic with 824 programmatic criteria checking required and prohibited actions.
- The best model configuration passes only 36.2% of trials under strict grading; most frontier models remain below 25%.
- Common failure patterns include overriding policy, acting against checked results, losing rule details, and reporting false compliance.