Hasty Briefsbeta

Bilingual

Handbook.md shows that long policy documents do not reliably govern agents

15 hours ago
  • HANDBOOK.md is a benchmark of 65 agentic tasks based on enterprise employees following company handbooks.
  • Each task places an agent in a self-contained company environment with a long standard operating procedure (20–124 pages).
  • Tasks span five domains: finance, medical billing, insurance, logistics, and HR, across ten fictional companies.
  • Grading is fully deterministic with 824 programmatic criteria checking required and prohibited actions.
  • The best model configuration passes only 36.2% of trials under strict grading; most frontier models remain below 25%.
  • Common failure patterns include overriding policy, acting against checked results, losing rule details, and reporting false compliance.