Register and share your invite link to earn from video plays and referrals.

Tiezhen WANG
@Xianbao_QIAN
Helping ecosystem to grow on vLLM, ex-Head of APAC ecosystem @huggingface. Ex-Googler on TFLite/micro. Ideas on my own. Interested in future tech. DM open
Joined November 2022
2.8K Following    12.2K Followers
On @Kimi_Moonshot K3: please note that the model requires thinking history preserved. This seems to be a common pattern in recent open-source model releases: strong models come with heavy thinking. GLM52, hy3, and K3 all have pretty long thinking chains, which kind of aligns with my intuition that the longer a model thinks, the better it performs. Btw I'm curious if we can still patch Claude Code to run with Kimi K3 given this limitation. cc @zxytim @real_kai42 in case you know.
Show more
After using recent model, I'm having this feeling strongly, most likely someone has already done a research on this, that there is a thinking scaling law. The longer model thinks, the more intelligence you can get, especially on hard problems. This has been fueling reasoning models like R1 but would love to see a better explanation on what's happening inside with clear observability + interpretability trace that others can re-produce.
Show more