For problem 1, you need to first see the journey of a weight during weight loading in vLLM. The journey is long and arduous:
1. Fuse
2. Relayout
3. Split
4. Shard
5. Copy into Buffer
6. Process
7. Copy into CUDA graph-captured memory
Operations 1-5 happen in the layerwise reloading stage in vLLM. Step 6 can involve a bunch of custom transformations like quantization, kernel format padding/striding, etc. We want to leave Step 6 to the vLLM engine and focus on steps 1-5 on the trainer.