Hasty Briefsbeta

Bilingual

RLVR might be disproportionately bad at science

17 hours ago
  • RLVR is ineffective for scientific discovery because scientific verification loops can span decades or centuries, and better theories often make worse initial predictions.
  • Historical examples like Copernicus vs. Ptolemy and the search for Neptune vs. Vulcan show that ex ante judgment of progressive vs. regressive research programs is nearly impossible.
  • Scientific breakthroughs require idiosyncratic biases, long-term dedication, and multiple concurrent research agendas, which RL loops cannot easily replicate.
  • Parallel discoveries (e.g., Darwin and Wallace) suggest that scientific progress depends on cumulative, circumstantial evidence and ancillary intuitions, not just verifiable predictions.
  • RL fails because science involves editing the objective (reward function) itself, not just optimizing within a fixed one—violating principles like Non-Circularity and Epistemic Sovereignty.