My first benchmark for vals, inspired by the model discovery agent work of
@sirbayes, testing mathematical intuition and experiment selection.
Most discriminative evals are long horizon — MysteryMechanism is unique in that short tasks separate the frontier very well.