가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Matej Sirovatka
@m_sirovatka
head of hr @ prime intellect | int64 upcaster
가입 August 2021
538 팔로잉 중    3.9K 팬
I'm not sure this is worded perfectly, but I agree. I think there is place for 3 types of inference engines for 3 separate groups: 1. Toy Inference engine - someone who wants to learn 2. Generic Inference engine - such as VLLM, SGLANG, etc - good performance out of the box, sensible API, etc - this is basically for 99.9% people 3. Custom Inference engine - if you want to squeeze out every bit of perf, have narrow requirements, need a different api plane, then you just write your own - you need a huge amount of resources to make something worthwhile over the current OSS defaults which are quite amazing
더 보기
my hot take is specialized inference engines are amateur projects to spend the spare tokens when @thsottiaux gives people one more reset😄 an inference engine is way more than just running a model on a certain hardware, it is an inference ecosystem. I use this slide several times when I talk to people about vLLM. Models, hardwares, and inference techniques, all of them move quickly. And vLLM lies in the intersection to provide a unified interface to end-users and applications. vLLM is the inference ecosystem. vLLM not only support current models and hardwares, but we are also working to support new models and hardwares coming in the next few months. It's an ecosystem people can trust and rely on. Over the past 3 years, I have seen so many projects claiming to be better than vLLM in certain aspects, but in the end either their techniques are contributed to vLLM or they disappear. That's the power of ecosystem. A specific example would be tilert , a megakernel inference engine dedicated for decode. They collaborate with vLLM by using vLLM prefill + tilert decode in , as highlighted in from @SemiAnalysis_ .
더 보기