First multi-turn agent harness RL training with sandboxing running at scale on the Hub ๐ ๐ค
Spawned 9,523 sandboxes in 14h.
Qwen3-Coder-30B-A3B (30B MoE) trained in full, FSDP2 + expert parallel.
All on HF Jobs. No Slurm anywhere.
0 crashes ๐ค