Hasty Briefsbeta

双语

Post-Training RLM Agents for End-to-End M&A Diligence

20 days ago
  • A recursive language model (RLM) harness improves rubric criteria pass rate by 39.1 percentage points on average across seven models for M&A diligence tasks.
  • Post-training the root agent within the RLM harness via reinforcement learning increased pass rate from 29.9% to 63.0% on held-out data rooms.
  • The root agent's coordination role is critical, as changing the root model had a larger effect on performance than changing sub-agents.
  • Self-distillation SFT improved pass rate from 46.1% to 60.1% by stabilizing effective delegation and memo writing behaviors.
  • RL training led to higher data room coverage (from 62% to 96%) and more sub-agent calls, without explicit coverage reward.
  • Ongoing work includes scaling RL with GLM-5.3 as the root model and exploring deeper recursion and joint training of sub-agents.