Dream-RSI: Recursive Self-Improvement through Evolving Worlds
3 hours ago
- Recursive self-improvement is crucial for autonomous AI agents but is bottlenecked by effective exploration strategies.
- Current systems struggle with fixed strategies that fail to adapt and expensive online policy optimization over large meta-search spaces.
- Dream-RSI introduces a lightweight orchestration layer that makes exploration explicit and programmable without altering the underlying coding agent.
- It uses a replay simulator built from historical discovery trees to provide immediate, low-cost off-policy feedback for refining exploration policies.
- Improved policies are redeployed online to drive further discovery, forming a self-improving loop that continuously expands the simulator pool.
- Dream-RSI achieves competitive or improved discovery quality across algorithm engineering, mathematical optimization, and GPU kernel engineering, with reduced discovery costs.