Fixing GRPO's credit assignment problem without evaluating every step
5 hours ago
- GRPO assigns uniform trajectory-level advantages to all policy tokens, failing to distinguish consequential decisions from less relevant ones.
- ProVer uses an agentic judge to contrast successful and failed trajectories and propose a potentially responsible segment.
- ProVer verifies the proposed segment by estimating advantage from the difference in terminal success rates before and after the segment.
- Positive advantage estimates are incorporated into GRPO advantages for tokens within the proposed segment.
- ProVer achieves 9.91% and 7.12% relative improvements over GRPO for Qwen3.5-2B and Qwen3.5-4B, respectively, on ALFWorld, WebShop, and SearchQA.
- Informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model.