가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Sumanth Hegde
@sumanthrh
Post-training @anyscalecompute. Prev - @UCSanDiego, @C3_AI, @iitmadras. Machine Learning and Systems. Intensity is all you need.
가입 February 2016
26 팔로잉 중    1.2K
And now for the second problem - scaling to large models. Here, we applied some tried and tested ideas: 1. PP-local and EP-local gather: Avoid redundant gather across PP groups, avoid gathering EP layers to avoid OOMs 2. Pipelined execution: Pipeline all-gather on the trainer, the weight transfer and the post-process on the inference side
더 보기