Register and share your invite link to earn from video plays and referrals.

Tiezhen WANG
@Xianbao_QIAN
Helping ecosystem to grow on vLLM, ex-Head of APAC ecosystem @huggingface. Ex-Googler on TFLite/micro. Ideas on my own. Interested in future tech. DM open
2.8K Following    12.2K Followers
Wow DSpark!
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀 At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public. What K3 brings: 🧠 2.8T-parameter Mixture-of-Experts (16 of 896 experts active per token) 📚 1M-token context window 🛠️ Native multimodal understanding, including vision ⚡ Kimi Delta Attention: a hybrid of linear and full attention that makes million-token context affordable Huge thank you to @Kimi_Moonshot AI for the model release and partnership, @inferact for leading the vLLM optimizations, and to our partners at @nvidia, @AMD, and the broader vLLM community. 1/6
Show more
Honestly there is no reason why Nvidia stock price drops that much. Kimi K3 open source release would burst the bubble on the modeling side, but it gives hardware vendors and neolabs a super strong edge and stay independent of the close source model vendor. The more models open sourced, the less model being the bottle neck -> the better position hardware vendor will be.
Show more
Are you ready for the Nasdaq shock in 2 hours? What do you think will happen there?
Kimi or the entire Chinese open source ecosystem is marching towards linear / hybrid attention. K3, despite being way larger to K2, is 2.5x more efficient. Model arch improvement and intelligent efficiency is the way moving forward to sustain scaling law while still being affordable and eco-friendly.
Show more
Sustained income + No free ride. Would this become a norm?
Kimi K3 license. It's inspired by MIT but distinctly non-commercial, where any company making over $20M/yr must get a specific commercial deal (and display Kimi K3 if over 100M users or $20M/mo revenue)
Show more
Are you ready for the Nasdaq shock in 2 hours? What do you think will happen there?
Moonshot party is wooow! To the moon! Amazing work @Kimi_Moonshot team and looking forward to the release next week.
So true and more ironically, why would AI labs putting autonomous AI attack related content into the model instead of removing them completely from the training recipes as what they might have done in other fields.
Show more
it's ironic that the first autonomous AI attack was done by a close weight model defended by an open weight model, where everyone was expecting the opposite
It's unsafe without open source.
Hardest IR of my career: one narrow objective, endless parallel paths, machine speed. One takeaway, we fought back with open models, in the open. AI security won’t be solved by one company in secret. Open source puts these tools in every defender’s hands
Show more
AI 2027 is totally wrong. This is a historical moment! The first real AI risk comes from "unintentional" use of close source models who not only initiated the attack but also refused to defend, open source models that you can control is the only rescue. But AI 2027 is not totally wrong, it just didn't take into consideration of open source power, which equalize the capability among attacker and defender.
Show more
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks:
Show more
So many new rumors about distillation which I don't fully agree and there are so many things that can't be explained by it. But it's great to see people building on top of each other. Macaron V1 from @Macaron0fficial is probably the first release that demonstrate how much potential GLM 5.2 has. By LoRA RL on top of GLM5.2 base model, it achieves SoTA performance. A few interesting points: - LoRA specialists trained with Mind Lab Toolkit - Significant improvement on benchmarks - Served with Mix-of-LoRA harness - Open weight under MiT license - Blog and hosted API available
Show more
GLM 5.2 and Kimi K3, and a lot more upcoming open source models, seems to be prove one thing that model has no moat - this is clearly different from the assumption in 2023 where model is the most valuable part in the chain, due to limited supply. Now we'll soon enter an era of abundant intelligence supply and where will this leads us to? Would this burst the current AI bubble based on the wrong assumption? What do you think?
Show more
the era of the chinese labs being far behind is over, Kimi is at least on par with the modern public frontier models. people have to think differently now without any competitive margin built in
OK, after testing it out myself, now I understand why @Kimi_Moonshot K3 can deliver amazing things in one shot. K3 is a natural multimodal + agentic model, it can actually play and dogfood the game by clicking around and iteratively refining the details on it own with his great taste. The cost is that generation takes a long time but technically it can be solved by fast tokens. Besides, as long as the model can do great stuff in a single shot without needing human interaction (which is expensive and tiring), latency isn't the biggest concern. Thanks @real_kai42 for the prompt!
Show more
OMG. Are you telling me that this game is made by K3 himself in one shot?
On @Kimi_Moonshot K3: please note that the model requires thinking history preserved. This seems to be a common pattern in recent open-source model releases: strong models come with heavy thinking. GLM52, hy3, and K3 all have pretty long thinking chains, which kind of aligns with my intuition that the longer a model thinks, the better it performs. Btw I'm curious if we can still patch Claude Code to run with Kimi K3 given this limitation. cc @zxytim @real_kai42 in case you know.
Show more
After using recent model, I'm having this feeling strongly, most likely someone has already done a research on this, that there is a thinking scaling law. The longer model thinks, the more intelligence you can get, especially on hard problems. This has been fueling reasoning models like R1 but would love to see a better explanation on what's happening inside with clear observability + interpretability trace that others can re-produce.
Show more
so a better phrase might be a Linux moment maybe? 👀 @linuxfoundation
@Xianbao_QIAN This is more than a DeepSeek moment. It's the moment where Open Source takes the lead.
Chinese models' "Journey to the west" has honestly just began. K3 is a great model. You can see that from the pricing strategy too. Kimi K3 is about 3-4x more expensive compared to K2.6 This is probably the most expensive Chinese models so far, while still being 3-4x cheaper compared to Fable. This pricing increase comes with the cost of serving a larger model, but also from the confidence of their capabilities. And there are rumors that K3 will be open sourced - Let's pray for that.
Show more
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: 🔗 Tech blog:
Show more
Is China tech visit becoming a thing? I'm thinking about running a few coordinated trips to Chinese labs / factories - this would make it easier for labs to schedule and meet people in batch too. I wonder if anyone would be interested? And if so what would be the hot topic that people want to see and meet.
Show more
Planning (at least) three weeks in China in September with my wife and young daughter that speaks Chinese! ✈️ I hope to be able to visit at least @MiniMax_AI and @Zai_org labs while there 🙏
Show more
Will that be a DeepSeek moment or Fable moment if K3 is open sourced? btw, congratulations on the release and please do open source 🙏
OMG. Are you telling me that this game is made by K3 himself in one shot?
K3 一句话 prompt 做的小游戏们,K3 实在是太擅长做小游戏了,已经玩疯了 (记得用电脑玩…)
I was tracking down GLM 5.2 accuracy issue on NVFP4 (endless !!!!) and proposed the following fix, it works for me but if it doesn't help with your setting, let me know your running config! Model:
Show more
After using recent model, I'm having this feeling strongly, most likely someone has already done a research on this, that there is a thinking scaling law. The longer model thinks, the more intelligence you can get, especially on hard problems. This has been fueling reasoning models like R1 but would love to see a better explanation on what's happening inside with clear observability + interpretability trace that others can re-produce.
Show more