๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Dirhousssi Amine
@DirhousssiAmine
๐Ÿ‡ฒ๐Ÿ‡ฆ ML engineer - post training team @huggingface ๐Ÿค— Rustacean ๐Ÿฆ€ โ— BJJ competitor โ— Lifelong Martial Artist
๊ฐ€์ž… March 2013
444 ํŒ”๋กœ์ž‰ ์ค‘    643 ํŒฌ
We ran async GRPO across HF Jobs with the trainer and vLLM on separate machines, no NCCL, and no shared disk. The adapter is mounted at the same path in every Job. LoRA training has landed in TRL's AsyncGRPOTrainer ๐ŸŽ‰๐ŸŽ‰๐ŸŽ‰ The trainer now syncs only the adapter to vLLM ~few MB for a 1.5B model. That changes what the setup can look like. We ran async GRPO across HF Jobs with the trainer and vLLM on separate machines, no NCCL, no shared disk. The adapter travels through a ๐Ÿชฃ mounted on the same path in every Job. Blogpost ๐Ÿงต
๋” ๋ณด๊ธฐ