Register and share your invite link to earn from video plays and referrals.

Search results for 涮乃葉
涮乃葉 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 涮乃葉
🔥 @TencentHunyuan Hy4 preview just cited a Zhihu post with a surprisingly simple challenge to DeepSeek’s mHC: What if the doubly stochastic matrix is overcomplicating things—and Identity actually works better? Today, let’s revisit the cited post, “Your DeepSeek mHC May Not Need the ‘m’,” by Zhihu contributor 涮月亮的谪仙人. After training Qwen3 1.7B and 8B dense models from scratch on 150B tokens, the experiments found: 💡 Identity HC > mHC > mHC-lite > orthogonal mHC In other words: simply setting H_res = Identity beat the Sinkhorn-Knopp constrained version. Here’s why 👇 1️⃣ What does mHC actually change? Standard Transformers have one residual stream. Hyper-Connections (HC) expand it to multiple parallel streams: 🔹 H_pre reads from the streams 🔹 H_post writes back to them 🔹 H_res mixes information between them DeepSeek’s mHC constrains H_res to a doubly stochastic matrix via Sinkhorn-Knopp, helping preserve norms and stabilize propagation. But the experiments suggest a simpler question: Do we need H_res mixing at all? 2️⃣ mHC seems to learn something close to Identity anyway For a single layer, the learned H_res is already close to Identity: diagonal ≈ 0.96, off-diagonal ≈ 0.01 But multiply H_res across many layers, and it gradually collapses toward a uniform 0.25 matrix. So each layer may look almost like Identity, while their cumulative effect becomes uniform mixing. The simplest fix? Just set H_res = I. 3️⃣ Identity preserves stream semantics With Identity, each residual stream stays where it is: Stream 0 stays Stream 0. Stream 1 stays Stream 1. No repeated reshuffling, no cumulative mixing, and Iᴸ = I. H_pre and H_post also no longer need to track where each stream has been repeatedly moved—they simply learn where to read and where to write. 4️⃣ But cross-stream communication still happens Setting H_res = I does not isolate the streams. The projection that generates H_pre and H_post already sees all residual streams, making H_pre input-dependent. So information can still be dynamically aggregated across streams before Attention/MLP and written back afterward. H_res isn’t the only mechanism for cross-stream interaction. 5️⃣ Why might Sinkhorn hurt? Repeated products of positive doubly stochastic matrices tend toward uniform mixing. In the Qwen3-1.7B experiment, after 56 HC modules, the minimum singular value of the accumulated H_res product reached just: 9.2 × 10⁻¹⁸ By ~10 layers, the four streams were already approaching the same 0.25 uniform mixture. Sinkhorn also comes with extra cost: 20 iterations, backward recomputation, extra parameters, and approximation error. Identity has none of these—and preserves the residual signal exactly. 6️⃣ More sophisticated alternatives didn’t win either The experiments also tested mHC-lite, softmax-weighted convex combinations, and orthogonal variants using Cayley/Givens transforms. The observed ranking remained: Identity HC > mHC > mHC-lite > orthogonal mHC The simplest design won. 💡 The takeaway DeepSeek’s mHC uses sophisticated manifold constraints to stabilize Hyper-Connections. But these experiments suggest that H_res itself may not need to be learned or mixed at all. Sometimes the best manifold constraint is the most boring one: H_res = I. Or, as the original post puts it: Maybe DeepSeek’s mHC doesn’t need the “m.” 😆 👉 Read the full Zhihu post for the training curves, mathematical analysis, H_res visualizations, and implementation details: #DeepSeek# #mHC# #Hunyuan# #Tencent# #LLM# #Transformer# #AIResearch# #AI#
Show more