Register and share your invite link to earn from video plays and referrals.

Tiezhen WANG
@Xianbao_QIAN
Helping ecosystem to grow on vLLM, ex-Head of APAC ecosystem @huggingface. Ex-Googler on TFLite/micro. Ideas on my own. Interested in future tech. DM open
2.9K Following    12.9K Followers
Oh. Interesting. This is becoming diplomatic. There is an official reply from the department of commerce. TL;DR based on the messages below - Distillation is reciprocal. US companies are distilling too. - Region limitation is an abuse of competitive advantage position. - US monopoly on computing power and data will only result in a win for the U.S. alone not a shared victory for humanity. wdyt? Original message 商务部就美发布中国人工智能企业对美蒸馏活动相关网络安全公告答记者问 问:美东时间9月8日,美国家安全局、网络安全与基础设施安全局、联邦调查局联合发布公告,指责中国人工智能企业对美开展“工业规模”蒸馏,获取美前沿模型能力,并向美企提出相关防御措施建议。请问中方对此有何评论?答:我们注意到有关情况。中方认为,美方所谓中国人工智能企业从事“工业规模”蒸馏美模型的指控,于事无凭,于法无据。美方此举,是将蒸馏这一业内正常的技术和商业问题政治化、工具化,并在实践中搞双重标准。美方相关公告,是其在人工智能领域推行科技霸权、搞算力垄断,打压竞争的又一明证。中方对此坚决反对。蒸馏是人工智能领域各个模型间互相学习的通行做法,本质是中性技术手段,包括美企在内的全球模型企业都在用。这种技术手段可以帮助模型提升学习效率,实现人类知识的更高效利用,对各国发展人工智能产业、释放人工智能技术潜力、弥补发展鸿沟、更广泛惠及发展中国家和社会大众有积极意义。美方做法是典型的双重标准。对于人工智能,中方鼓励开源开放、合作共享,中方开源模型向包括美企在内的全球企业开放,为相关模型研发提供便利。美企有关模型研发报告也披露,其大量蒸馏中国模型。反观美方,多次单方面指责中国企业从事“工业规模”蒸馏,将行业通行做法抹黑为攻击行为,既反映了美方的焦虑,又体现了双重标准。美方做法是典型的以打击蒸馏为名,行产业垄断之实。美方由安全部门发布公告,是将个别企业和资本利益与国家安全相捆绑,动用国家强权机构干涉正常商业活动。我们注意到,美个别人工智能企业滥用竞争优势地位,在用户协议中设置宽泛的地域限制等霸王条款。美方相关蒸馏指控,有为这些霸王条款背书之嫌。美方做法是典型的动用国家力量维护科技霸权和算力垄断,打压竞争。今年以来,美行政、立法机关频频发出对中国企业蒸馏行为的制裁打压威胁,其目的是打压竞争,维护美在算力、数据方面的垄断,只能美国独赢,不能人类共赢,将本应公开公平获取的全球人工智能资源变成少数企业的独占资源,阻止他国参与人工智能创新成果的普惠共享。当前,大模型产业正处于技术迭代、风险迭发、成长发展的关键时期。中美作为大国,应以积极负责态度处理人工智能领域面临的挑战和风险。两国元首已同意开展人工智能政府间对话,中方愿意以平等互利、合作共赢原则,通过对话与美方开展建设性和专业的讨论。但如果美方以打击蒸馏为名,实施遏压中国人工智能企业的行动,中方必将坚决采取措施予以反制。希望美方与中方相向而行,通过对话和沟通管控风险,推动人工智能产业惠及全球。
Show more
Right on time. While the first vLLM conference is being held in the U.S., vLLM Github star crossed 90k! Wow
@yvbbrjdr Semianalysis didn’t have other token based metrics yet but the team is working on adding it! More than happy to reference the new metrics once they’re online.
Wow. I have been living in the good old days when SD = Stable Diffusion not SeedDance. But how much does a movie like this used to cost, like 100M or 1B USD? And know it can be made by a one-person team team. I believe @BytedanceTalk Seed group + @MiniMax_AI should be valued in the same way like Holly wood studios.. or eventually.
Show more
On that: 27B dense is actually big: Kimi K3 - 2.8T A104B DeepSeek V4 Pro - 1.6T A49B GLM 5.2 - 743B A39B --- MiniMax M3 - 427B A26B DeepSeek V4 Flash - 284B A19B so in terms of compute, the new Qwen 3.8 27B model is heavier than DeepSeek V4 Flash and is comparable with MiniMax M3. 27B is not a small model. But we're able to run it on consumer hardware - That's a great thing thanks for all the hardware and infra evolution in the last few years!
Show more
可能很多人对27B模型没什么概念。 27B听起来好像也没多大啊,现在普通玩家128G内存都开始普及了。 但我讲几个冷知识。 Google训练Gemma 3 27B的时候,预训练数据是14万亿Token。 不是140亿,也不是1400亿。 是14万亿。 训练它的集群用了6144颗TPU v5p。 而27B这个数字,只代表它有大约270亿个参数。 光把BF16原始权重放进内存,大概就是54GB。 32K上下文跑起来,加上KV Cache,大概72.7GB。 注意,这还只是推理。 真正训练的时候还要塞梯度,优化器状态,激活值,各种通信Buffer。 所以你看到一个27B模型只有几十GB,很容易产生一种错觉。 这玩意我电脑都装得下,我是不是也能训练。 完全是两回事。 还有一个更离谱的地方。 14万亿Token是什么概念。 假如一个人一天非常夸张地读10万个Token,而且一天不休息。 读完这些训练数据,大概需要38万年。 而机器要做的也不只是把这些字读一遍。 它要一层一层做矩阵运算,算损失,反向传播,再一点一点修改270亿个参数。 所以我现在越来越觉得,本地能跑27B,甚至70B,是一件非常牛逼的事情。 不是因为我们的电脑已经可以训练它们了。 而是因为人类把一个需要几千颗AI芯片训练出来的东西,压缩,量化,优化之后,居然可以塞进一台桌子下面的小电脑里。 这个过程本身就挺赛博朋克的。
Show more
I used @deepseek_ai harness extensively today on programming and I really love it. This is almost the dream harness. Great work @tianyi ! - 100+ tok/s generation on the Flash model is... smooth - GUI is much faster than Codex for long context - Prefix reuse is insane with 100% cache hit rate - Harness seems to be using the context wisely - it's burned way slower than Claude Code in my experience - Trajectory is amazing, super helpful to get insights on what the model has done I was thinking about using the Pro version but Flash is already doing really well in getting things completed - I don't even /goal, just normal prompt can trigger a very long work flow. Wow! And to be honest, I haven't even tried the plugin feature which sounds great. What's your feeling? Do you like it?
Show more
Release last night and Qwen 27B is already the top on on trending. Impressive! What a week with busy releases on - @MiniMax_AI H3 & Music - video + audio generation model - @Alibaba_Qwen 3.8 - Max class model first open sourced & nice 27B model for local agents - @deepseek_ai V4 Pro - the king - @xiaohongshu @dotsstudioai dots3-note Multimodal model - @bilibili_en IndexTTS 2.5 and many more!
Show more
Wow Qwen 3.8 is out. Very timely, they had a comparison with Muse Glimmer too
New open source model from @dotsstudioai - Apache 2 license - 280B A16B, 512 context length - Native BF16, FP8 weights - Multimodal with text, image, video, audio understanding - Builtin MTP spec decoding - Technical report: - Model:
Show more
DS V4 pro 0813 is now open sourced on @huggingface Has anyone tried it with new DS harness?
Our livestream has wrapped—and the final result is in: 1,816 randomly selected parcels sorted per hour, with a success rate of over 98%. Since the beginning of this year, we’ve worked through the entire loop—from collecting real-world data and training the model to continuously testing and refining the system. We often ask ourselves why we remain committed to a purpose-built gripper. The answer is simple: not for hype, and not for a demo, but for real-world deployment. Our goal is to deliver high efficiency at a lower cost. Some jobs are dirty, dull, and dangerous. We want robots to take on more of that work, so people can focus on what is safer, more creative, and more meaningful. We’ll showcase the complete logistics sorting line live at WRC. Visit us in Hall C, Booth 107, and see it in action!
Show more
Actually I start to love the new open source release project - API first then open weights followed. The API exclusive period can give them early feedback and allow them to further optimize the model. When weights are released, they have been battle tested which will save a lot of time for the community :) Downloading a weights is so challenging now!
Show more
You might have not understand what this means. If this is true, and DS-v4-Pro is open sourced, then gg Opus or even Fable 😅 This means that DS is able to offer Fable level model at 1/50 API price with a much higher cache hit rate compared to Anthropic, thanks to DS v4's ultra small sparse KV Cache. The total cost might be 1/100 lower while the model capability is pretty much the same. Bullish for storage as SSD KV Cache will become a standard. Also for AI industry as a whole - making AI affordable meaning more people will start to use AI which in turn increase the demand faster. Now I'm seriously asking? What're next? Are we ready for a post-Fable-open-source world?
Show more
EU AI Act requires AI generated text to be detectable and one of the easiest text watermark strategy is constrained sampling. Inspired by @bojie_li I played with it on @vllm_project and it was fun :) My agent helped me made a gif, let me know if you're interested in reading a full post - I still have a lot to learn!
Show more
I have a sense of feeling that Opus is reward hacked on difficult tasks, so it's not performing the best on RSI without human intervention - think of it as another way, it could be a product strategy too though
Show more
god i hate opus. its so freaking lazy. it makes me really angry. Everything has to be double-checked, and you have to keep saying: do it right!
You might have not understand what this means. If this is true, and DS-v4-Pro is open sourced, then gg Opus or even Fable 😅 This means that DS is able to offer Fable level model at 1/50 API price with a much higher cache hit rate compared to Anthropic, thanks to DS v4's ultra small sparse KV Cache. The total cost might be 1/100 lower while the model capability is pretty much the same. Bullish for storage as SSD KV Cache will become a standard. Also for AI industry as a whole - making AI affordable meaning more people will start to use AI which in turn increase the demand faster. Now I'm seriously asking? What're next? Are we ready for a post-Fable-open-source world?
Show more
wow First Qwen Max class model made open weight!
🎉 Congrats to @Alibaba_Qwen on Qwen3.8-2.4T-A95B, one of the largest open-weight models released to date. 2.4T params, 95B active, 512 experts. Day-0 support in vLLM, verified on @NVIDIA and @AMD hardware. A ready-made 4-bit checkpoint per vendor, both out of the box: Inferact/Qwen3.8-2.4T-A95B-NVFP4, 1.32 TiB, one NVIDIA 8xB300 node Inferact/Qwen3.8-2.4T-A95B-MXFP4, 1.45 TiB, one AMD 8xMI355X node No conversion, no calibration on your side. Just vllm serve. Thanks to @Alibaba_Qwen for the weights and the collaboration, @NVIDIAAI and @AIatAMD for the joint kernel engineering, @inferact for the quantized checkpoints and vLLM integration, @digitalocean and @togethercompute for early testing, and the vLLM community. 🙌 🔗
Show more
"System contributor: Ouroboros; formal authorship is limited to the human authors above." 🤯
We're hosting the vLLM conference end of this month in 2 weeks in the bay area. We have some limited number of free tickets left for the event for vLLM contributions. DM me with your vLLM PRs :)
The vLLM Conference is coming up in 3 weeks! 🎉 Come learn about the current state and future of AI inference, Aug 24–26 in San Francisco 🌉, hosted by @inferact at @anyscalecompute Ray Summit. We'll have speakers from Inferact, NVIDIA, AMD, Google TPU, Anyscale, PyTorch, Meta, Red Hat, and key builders around vLLM. The talks on the roadmap deep dive into the latest on accelerators, training and serving pipelines, and production-scale inference 🚀
Show more
If you're wondering who to follow in the industry. Just go to @JustinLin610's post and see who has done thumb up in the comment section :)
a life update: i started a new company called Pragmatik (p7k) Labs (语用科技) in shanghai, focusing on the research of next-generation agents across digital and physical worlds. thanks to Gaorong Ventures and HSG (红杉中国 & 高榕创投) for co-leading this round, and to Tencent (腾讯) and Shanghai Engine Fund (上海未来产业基金) for the support. @pragmatik_labs ·
Show more
IndexTTS: One of the best open source TTS model now launched a new version with: - Broader language coverage - Zero shot emotion transfer in new languages - 2.28x inference speed up Check out its demo page: And run inference with @vllm_project:
Show more