Register and share your invite link to earn from video plays and referrals.

Adithya S K
@adithya_s_k
Scaling RL Envs @huggingface ๐Ÿค— โ€ข Founded @cognitivelab_ai Prev : Research @MSFTResearch โ€ข ML @apple โ€ข 22
Joined June 2020
2.4K Following    15.1K Followers
You can now do RL weight sync over HF Buckets. > Async GRPO + LoRA in TRL > train on one Job, serve vLLM on others > sync just the small delta adapter through a Bucket. Trainer, inference and storage, all on @huggingface infra. Beautifully optimised loop by @DirhoussssiAmine ๐Ÿ‘‡ 3h27 โ†’ 53 min for the same reward.
Show more
We ran async GRPO across HF Jobs with the trainer and vLLM on separate machines, no NCCL, and no shared disk. The adapter is mounted at the same path in every Job. LoRA training has landed in TRL's AsyncGRPOTrainer ๐ŸŽ‰๐ŸŽ‰๐ŸŽ‰ The trainer now syncs only the adapter to vLLM ~few MB for a 1.5B model. That changes what the setup can look like. We ran async GRPO across HF Jobs with the trainer and vLLM on separate machines, no NCCL, no shared disk. The adapter travels through a ๐Ÿชฃ mounted on the same path in every Job. Blogpost ๐Ÿงต
Show more