19 hours ago
- RLVR is ineffective for scientific discovery because scientific verification loops can span decades or centuries, and better theories often make worse initial predictions.
- Historical examples like Copernicus vs. Ptolemy and the search for Neptune vs. Vulcan show that ex ante judgment of progressive vs. regressive research programs is nearly impossible.
- Scientific breakthroughs require idiosyncratic biases, long-term dedication, and multiple concurrent research agendas, which RL loops cannot easily replicate.
- Parallel discoveries (e.g., Darwin and Wallace) suggest that scientific progress depends on cumulative, circumstantial evidence and ancillary intuitions, not just verifiable predictions.
- RL fails because science involves editing the objective (reward function) itself, not just optimizing within a fixed one—violating principles like Non-Circularity and Epistemic Sovereignty.