Current models are so powerful that this problems gotten worse, because going in subtly wrong directions can be hard to detect! Y'day I realized that Astra had been running an eval on some random synthetic subset of the data instead of the bench, as I asked. Cost me a day!