~6-7 million tokens per game for the cheapest run is quite a lot. many many reasoning tokens. def trained directly on such an env (no reasoning vs low reasoning on the standard harness also interesting)
"contain" "it's worst" and "resemble our internal monitoring deployments" (which we recently demonstrated are keeping logs and not looking at them until after the fact) is not super reassuring.
so:
- any task/environment that you can score reliably and invest enough money (= an obscene amount) in bruteforcing its training, will be solved
- every challenging and popular benchmark will meet this bar
- abilities won't necessarily transfer
how can we measure progress?
please don't title your papers for "maximizing engagement" by removing key information (=lying to people). we will find out in a minute anybow, so why waste our time? make titles as informative as possible, so we can easily find it when searching, and easily vet it while reading.