On
@Kimi_Moonshot K3: please note that the model requires thinking history preserved.
This seems to be a common pattern in recent open-source model releases: strong models come with heavy thinking. GLM52, hy3, and K3 all have pretty long thinking chains, which kind of aligns with my intuition that the longer a model thinks, the better it performs.
Btw I'm curious if we can still patch Claude Code to run with Kimi K3 given this limitation.
cc
@zxytim @real_kai42 in case you know.
After using recent model, I'm having this feeling strongly, most likely someone has already done a research on this, that there is a thinking scaling law.
The longer model thinks, the more intelligence you can get, especially on hard problems.
This has been fueling reasoning models like R1 but would love to see a better explanation on what's happening inside with clear observability + interpretability trace that others can re-produce.
Show more