注册并分享邀请链接,可获得视频播放与邀请奖励。

Sumanth Hegde
@sumanthrh
Post-training @anyscalecompute. Prev - @UCSanDiego, @C3_AI, @iitmadras. Machine Learning and Systems. Intensity is all you need.
加入 February 2016
26 正在关注    1.2K 粉丝
For problem 1, you need to first see the journey of a weight during weight loading in vLLM. The journey is long and arduous: 1. Fuse 2. Relayout 3. Split 4. Shard 5. Copy into Buffer 6. Process 7. Copy into CUDA graph-captured memory Operations 1-5 happen in the layerwise reloading stage in vLLM. Step 6 can involve a bunch of custom transformations like quantization, kernel format padding/striding, etc. We want to leave Step 6 to the vLLM engine and focus on steps 1-5 on the trainer.
显示更多