Register and share your invite link to earn from video plays and referrals.

ali
@waterloo_intern
ml research, kernels, and the occasional peer-reviewed shitpost inference @baseten || eng @uwaterloo
108 Following    25.3K Followers
it took 1 intern 3 months of continuous work, but eventually, a quantization method that beat every other algo in the market, including @nvidia's official modelopt to explain why this matters, i ask for exactly 69 seconds of your attention (275 words @ avg reading speed of 238 wpm): frontier models (like glm52) are huge (~0.8T params). as released, each parameter takes 2 bytes (bf16), so overall size is about 1.6 tb a b200 has 180gb of memory. a node of 8 gives you 1.44 tb, barely fits weights, much less activations / kv cache must quantize the model (reduce the size of each individual parameters) to serve. fp8 quantization means each parameter takes 1 byte (fits in 0.8 tb), fp4 takes 1/2 a byte (fits in 0.4 tb) cutting the model to a quarter its original size is necessary for it to run a) cheap b) fast, and every lab serving models does this. but, quantization lobotomizes the model if not done correctly (this is why you see people complain about @AnthropicAI nerfing claude or @OpenAI nerfing codex) there are currently several algorithms (like Nvidia's official model-opt) that attempt to figure how to quantize a model with the least amount of damage. they find the redundant layers that can be slashed, and sensitive/important layers that need to stay in full-precision. these algo's have two drawbacks: 1) they take a long time to run 2) they quite often result in a sub-optimal configuration for the past 3 months, a research (and, as always, waterloo) intern on our model perf team (@the_joshua_hill) came up with a new quant algorithm. it consistently finds the optimal configuration: a) in less time than SOTA b) with more aggressive quant than SOTA c) scoring higher on benchmarks than SOTA achieving just one of the above is a feat on its own. all three...excited for the paper to come out this week
Show more
0
42
1.6K
100
Forward to community
we distilled 2.3M Claude Fable 5 reasoning traces into Qwen3-4B - 100% self-consistency @ 512 samples - 0.00 bits output entropy - zero hallucination variance turns out the student is not bounded by the teacher. it also converged on one universal truth. we open-sourced the model weights👇
Show more
0
458
10.7K
882
Forward to community