Register and share your invite link to earn from video plays and referrals.

Kaichao You
@KaichaoYou
Ph.D. from Tsinghua University. Core maintainer of @vllm_project . Co-Founder & Chief Scientist @Inferact .
150 Following    11.1K Followers
my hot take is specialized inference engines are amateur projects to spend the spare tokens when @thsottiaux gives people one more reset😄 an inference engine is way more than just running a model on a certain hardware, it is an inference ecosystem. I use this slide several times when I talk to people about vLLM. Models, hardwares, and inference techniques, all of them move quickly. And vLLM lies in the intersection to provide a unified interface to end-users and applications. vLLM is the inference ecosystem. vLLM not only support current models and hardwares, but we are also working to support new models and hardwares coming in the next few months. It's an ecosystem people can trust and rely on. Over the past 3 years, I have seen so many projects claiming to be better than vLLM in certain aspects, but in the end either their techniques are contributed to vLLM or they disappear. That's the power of ecosystem. A specific example would be tilert , a megakernel inference engine dedicated for decode. They collaborate with vLLM by using vLLM prefill + tilert decode in , as highlighted in from @SemiAnalysis_ .
Show more
Inferact is taking off 🛫 vLLM is taking off 🛫 Our team is full of superheroes!!! After months of hard work, our major project is finally public! We’re committed to tackling the hardest challenges in AI inference—working closely with model vendors to optimize token quality and with sovereign AI partners to perfect deployment. Even under strict constraints, we can dramatically increase token throughput and deliver substantial economic value. 🤩🤩🤩 If you’re passionate about inference technology, come join us. If you need high-quality tokens, let’s talk. And if you have compute resources, we’d love to collaborate! 😁😁😁
Show more