Register and share your invite link to earn from video plays and referrals.

熊师傅 weight decay 了吗
@bigeagle_xd
a person living with @mianmoe, opinions are my own
1.3K Following    11.9K Followers
I've decided it's time to resign from FAIR. 🫡 I'm especially grateful for the compute that made much of my research possible: 400 H100/200 GPUs allocated to me by FAIR; over a thousand H100s borrowed from FAIR Europe's CodeGen team led by Gabriel (@syhw ); thousands more borrowed from a legacy cluster in early 2025; and access to thousands of idle and low-prio GPUs across FAIR. For clarity, since joining FAIR four years ago, I've been on FAIR-level pay throughout. I mention this simply to avoid any confusion in any media coverage. Many thanks to @ylecun for founding FAIR, and to Joelle (@jpineau1 ) for steering FAIR through the later years of its golden age. じゃあね — so long, FAIR.
Show more
0
58
1.3K
42
Forward to community
Unreal views of #Zhuque3# Y2 first stage nailing its landing in Minqin today. It's such a clean return and recovery!
to my knowledge this is the largest open experiment on autonomous agents iterating on a research environment we scaled runtime, compute, diversity of models and harnesses. as a comparison, similar tasks on oai/anthropic system cards are anthropic "optimizing an llm training on CPU" and openai gpt 5.6 doing nanogpt track 1 but for less than a day. we also share a lot of the details (traces, scratchpads, ect..) so you can look into how models approach such tasks this experiment is quite noisy, one run in the same setting has a ~50 step spread after 24h. i find it super impressive that while models explore relatively the same ideas and the task and environment have a lot of variance, there is still a big gap between different models fable 5 closed 82% of the gap to the current human record with kimi K3 being very impressive as well. i'm currently running grok 4.6, deepseek v4 pro, muse spark 1.2, qwen 3.8 max and glm 5.3, expect results next week one of my favorite parts is some of the tools that models developed during this experiment, especially with prime agent. for instance kimi K3 created its own experiment API for generating optimizer variants, loss comparisons, Newton Schulz tuning etc.. we also have an early deepseek v4 pro <> prime agent run that did a PSGD experimentation outside the normal nanogpt loop to build intuition before starting gpu runs. other cool examples in the blog with other harnesses as well! i'm also very excited about the ideas we have in mind on this subject, we will keep working on understanding the research capabilities of (closed and open) frontier models
Show more
> for instance kimi K3 created its own experiment API for generating optimizer variants, loss comparisons, Newton Schulz tuning etc. awesome🕶️🕶️🕶️, speedrun was also what @nathancgy4 and I had in mind for a hill climbing showcase — didn't expect K3 to be this good at it. Kudos!
Show more
A hidden detail in the recently released Muse Glimmer model: its per-matrix weight RMS norm is pinned almost exactly at ~6e-3! If you’re curious why fixing weight RMS like this can work, feel free to checkout Hyperball:
Show more
Songlin Yang's video explanations Songlin Yang's blog post Design intuition Kernel algorithm (I skipped to "A Chunkwise Algorithm for DeltaNet" section) FlashKDA I prompted Codex to explain it to me, only after that the deep dive made sense vLLM serving explanations The great Zhihu post for explaining AttnRes Training Inference Jianlin Su's MoE 環遊記 series Explains the concept of MoE from math first principles Papers Kimi Linear, LatentMoE, Kimi K3 tech report N/
Show more
New RSIBench-Data ( result: Kimi K3 + Kimi Code achieves a 27.317% weighted score (aime,gpqa use 0.1, others 0.2) across six benchmarks, ranking #1# among the five model–harness combinations we tested. It scores 50% on SWE-bench Verified and 17% on SWE-bench Pro. A strong result for automated RSI research. Congratulate!
Show more
Deconstructing Scaling Laws: The Triad of Optimization, Architecture, and Data
fun fact: almost all diagrams in the report are made with TikZ written/edited by K3, who has demonstrated remarkable multi-format programmatic rendering capabilities. btw, as far as i know, k3 is the 2nd best TikZ master after @yzhang_cs 😉
Show more
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
Show more
a great summary of the past year's work. time to move on and carry on 🫡
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
Show more
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: Tech report: Tech blog:
Show more
0
1.5K
46K
7.2K
Forward to community
Understanding Reasoning from Pretraining to Post-Training A controlled study using chess to show that pretraining loss predicts post-RL performance and that RL discovers new correct moves on hard puzzles.
Show more
we're working extremely hard with our partners to open the weights as soon as possible
0
42
1.7K
58
Forward to community
Fun fact: I'm responsible of editing the limitations part, and this final line was written by Zhilin himself.
developing an "agent platform" is just like developing an operating system