The Rearrangement Inequality and Its Generalizations
Starting from the rearrangement inequality, this article follows two lines of development — "matricization" and "multi-sequence extension" — to survey its principal generalizations.
Show more
Steepest Descent on Manifolds: 7. The Close-form Solution of Stiefel
What Loss Functions Can LMs Use Besides Cross-Entropy?
A Brief Look at K3's MoE and Attention
Deconstructing Scaling Laws: The Triad of Optimization, Architecture, and Data
arrived as promised
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights:
Tech report:
Tech blog:
Show more
Linearizing Softmax Attention into Gated DeltaNet
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API:
🔗 Tech blog:
Show more
Similarity Measure Based on the Rearrangement Inequality
Taylor Expansion of LogSumExp and Softmax
Revisiting Convergence Results in Convex Optimization (VII)
- New identity quantitatively link weight averaging and learning rate decay
- discuss why uniform averaging loses to sliding, and why Schedule-Free still can't fully escape scheduling
Show more
The Beauty of Brute Force: A General Matrix Function Approximation Framework
Margin-Enforcing Projection
MoE (9): The Gate Normalization Debate
Steepest Descent on Manifolds: 6. Muon + Double Rotation
Introduces MuonR — a Muon variant that constrains updates to left & right rotation matrices. This preserves the singular value distribution of weights, providing a clean, elegant way to maintain training stability.
Show more
Why Does KellerJordan's Muon Have an Extra max(1, ⋅) Compared to MuP ?
Is Higher Singular Value Entropy Always Better for Matrix Parameters?
MoE (8): Enforcing Sequence-Level Balance
This article explores how to achieve sequence-level load balancing without incurring any loss penalty. Starting from the original Quantile Balancing (QB), we gradually derive a new method called Moving Quantile Balancing (MQB), which successfully accomplishes this objective.
Nevertheless, whether sequence-level balance is truly necessary and to what extent it should be enforced remain open questions.
Show more
How Does DeepSeek V4's tid2eid Come About?
This article briefly outlines the basic idea of Hash Routing in DeepSeek V4, with a particular focus on the construction principles of its tid2eid mapping table.
Show more
FID as Loss Directly: From Gradient Calculation to Stream Training