Hasty Briefsbeta

Bilingual

DoGBench: The first user-facing docs generation benchmark. No model scores >50%

4 hours ago
  • DoGBench is a benchmark for evaluating AI agents that write and maintain user-facing software documentation using real repository events and reported gaps.
  • It includes 292 tasks from open-source projects: 205 requiring documentation changes and 87 requiring no changes, with primary evaluation on a 117-item held-out split.
  • Agents struggle to decide when documentation needs updating, both over-updating and missing necessary changes.
  • The highest combined score among evaluated systems was 47.3/100, achieved by Qwen3.8 Max with OpenCode.
  • A broad audit found that 45.5% of submissions had task-completion gaps, 36.6% had technical inaccuracies, 32.5% omitted central concepts, and 6.1% contained fabricated content.
  • Expert review remains necessary; the highest P0-clean delivery rate was 39.0% (GPT-5.6 Sol with Codex).
  • Evaluation uses task-specific rubrics with over 3,000 criteria, and failing any P0 criterion caps a patch's score at 60 out of 100.
  • DoGBench separately reports update decisions (correct abstention) and patch quality, using the harmonic mean for the combined score.