Register and share your invite link to earn from video plays and referrals.

Sumanth Hegde
@sumanthrh
Post-training @anyscalecompute. Prev - @UCSanDiego, @C3_AI, @iitmadras. Machine Learning and Systems. Intensity is all you need.
26 Following    1.2K Followers
A higher KV cache hit rate doesn’t always mean lower TTFT and TPOT, or higher throughput. We studied LLM routers with different request scoring objectives under agentic and RL workloads. Here’s what we learned. 🧵
Show more
And this work is still early - this is a first step and demonstrates performant weight transfer with RDT. We are looking forward to improving the weight transfer implementation with some of the latest updates on the RDT APIs.
Show more
The blog has more details on the implementation as well as a demonstration of fault tolerance with RDT + NIXL. Check it out: Work led by @AaronHao4
We also studied how much of a difference different optimizations made for weight sync for Qwen3 235B on 4 H100 nodes.
Putting these together, the weight sync for an attention layer will look like follows:
And now for the second problem - scaling to large models. Here, we applied some tried and tested ideas: 1. PP-local and EP-local gather: Avoid redundant gather across PP groups, avoid gathering EP layers to avoid OOMs 2. Pipelined execution: Pipeline all-gather on the trainer, the weight transfer and the post-process on the inference side
Show more
The operations in weight loading can vary a LOT across models and layers. To support these, we use a “recording tensor” dry run - we use a dummy tensor to record the various transformations in the `load_weights` function, and store the op chain at init time. We replay the op chain on the trainer side to compute the full weights during each weight sync.
Show more
For problem 1, you need to first see the journey of a weight during weight loading in vLLM. The journey is long and arduous: 1. Fuse 2. Relayout 3. Split 4. Shard 5. Copy into Buffer 6. Process 7. Copy into CUDA graph-captured memory Operations 1-5 happen in the layerwise reloading stage in vLLM. Step 6 can involve a bunch of custom transformations like quantization, kernel format padding/striding, etc. We want to leave Step 6 to the vLLM engine and focus on steps 1-5 on the trainer.
Show more
There were two key problems to solve for sharded weight transfer: 1. How do you compute the shard to send to a vLLM worker on the trainer? 2. How do you scale to large models without blowing up memory and keeping transfer times low?
Show more
Excited to share our work on large-scale sharded weight transfer! We’ve implemented a native sharded weight transfer engine in vLLM and SkyRL using Ray Direct Transfer (RDT) + NIXL. We’re able to achieve weight transfer for Kimi K2 (1T params) in BF16 in 7.53 seconds on 48 8xH100 nodes. Blog:
Show more
Join us for the first-ever SkyRL Meetup! It's been a year since we released the SkyRL library. SkyRL now has 2k+ GitHub stars and a growing user base in academia and industry, including researchers from Stanford, CMU, UC Berkeley, Microsoft AI, Datadog, and more. We're hosting a SkyRL Meetup on August 18th at the Anyscale office in SF. We have a special set of lightning talks: 🔗 Making frontier RL training accessible to everyone, from @trajectorylabs 🐶 Post-training an SRE agent with SkyRL, from @datadoghq ⚡️ Scaling RL with SkyRL on AMD GPUs, from @AMD 🧑🏻‍💻 Training knowledge work agents with SkyRL, from the SkyRL team (@charlie_ruan) We’ll also do a deep dive into the library and talk about the road ahead. More details here:
Show more
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
0
563
14.9K
2K
Forward to community