Since I went into this release:
*It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo.
*Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm.
*Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.
one of the perks of working on ai is seeing things unroll very gradually: known about pangram for +2 years (almost met Max in NY while they were launching), and only now crashing on French twitter hard
Misaligned AI is the last thing India should worry about. We have misalignment of a billion people to start with as a problem. We have misaligned humans sitting in power we are not able to get rid of.
wait so what does "pacing the frontier" mean other than "we will voluntarily let the same third-party testers we work with for pre-release model evaluations continue to do pre-release model evaluations but like more and stuff"
+1 to my theory that the frontier labs are soon pivoting to automated research labs, bc Fable+ models are too expensive for most usecases but highly positive ROI for autonomous research