When you read a closed source blog pose, and they say "RL systems at scale take lots of engineer work", this is exactly the sort of thing they mean.
Really great work by the sand team to go through the hard science and engineering effort to get low precision RL working AND go through the additional effort to communicate it so clearly.
Take some time today to read through the blog, there's lots of good tidbits
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training