- Chip placement is crucial in physical design, but RL-based methods focusing on wirelength optimization often fail to achieve expert-quality layouts.
- The reward design is identified as the main cause of the performance gap with experts.
- The approach bypasses formalizing complex processes by learning directly from expert layouts to derive a reward model.
- It infers step-by-step expert trajectories from final expert layouts.