Intrinsic discovery uses RL to uncover new behavior in an environment. The model is rewarded for reaching states unlike those found so far, then sampled again after each update to push exploration further.
This approach discovers 5× more novel states than a frozen baseline.