At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training