注册并分享邀请链接,可获得视频播放与邀请奖励。

Eric Zhang
@ekzhang1
Computer systems person, interaction designer. @thinkymachines prev founding eng @modal → dreams of: a simpler, more honest, more human sort of software
加入 June 2021
651 正在关注    23K 粉丝
So… I have it working now. The core of an inference engine in Rust is about 3k lines (scheduling, kv/swa/recurrent state, prefix cache, multimodal, dynamic batching) I then asked Codex to implement (with MTP, cudagraphs, TP on Blackwell): - DeepSeekV4 Flash - Gemma 3.8 27B - GPT-OSS - Inkling It turns out that most models take a lot of code to run with fast, persistent graphable kernels and so on… on the order of 2.5k for a simple model (gpt-oss) and up to 7k for a monster like DeepSeekV4 Flash that has mHC, CSA, HCA and DSpark Not sure if I’ll continue with this, but it was fun! The simple core is reusable, and startup times are crazy fast for an inference engine with no tech debt and not having to serve a ton of dynamic configurations - just one model forward pass in nvfp4 :’) - maybe one day SGLang/vLLM will be modular like this. See interactivity-throughput below for Inkling-Small TP=1 after a few sconv kernel fusions. (This is without overlap scheduling btw! Don’t even really need it when your core + scheduler is small and Rust.)
显示更多
0
27
563
30
转发到社区