Reinforcement learning rewards convergence, and convergence is the opposite of discovery.
We built the world’s first judgment model from the decisions papers erase: what to pursue, challenge, abandon, and revisit. Columbus-1 autonomously found a zero-click Bluetooth RCE and designed a 10-foot, self-landing rocket.