We found a general recipe for Solves: scale pretraining, then hill-climb with minimal in-house data.
For the first time, one fine-tuning example can teach a new behavior that generalizes. Below: 4 folding strategies, learned from one example each and tested on held-out setups.