Hasty Briefsbeta

Bilingual

Fixing GRPO's credit assignment problem without evaluating every step

5 hours ago
  • GRPO assigns uniform trajectory-level advantages to all policy tokens, failing to distinguish consequential decisions from less relevant ones.
  • ProVer uses an agentic judge to contrast successful and failed trajectories and propose a potentially responsible segment.
  • ProVer verifies the proposed segment by estimating advantage from the difference in terminal success rates before and after the segment.
  • Positive advantage estimates are incorporated into GRPO advantages for tokens within the proposed segment.
  • ProVer achieves 9.91% and 7.12% relative improvements over GRPO for Qwen3.5-2B and Qwen3.5-4B, respectively, on ALFWorld, WebShop, and SearchQA.
  • Informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model.