Register and share your invite link to earn from video plays and referrals.

Search results for DELTARUNE
DELTARUNE community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including DELTARUNE
DELTARUNE Chapter 5 is out now!!!
0
1.4K
95.2K
16.1K
Forward to community
DELTARUNE tomorrow...
0
4.9K
254.9K
79.2K
Forward to community
DELTARUNE Chapter 5!!! Comes out in a little less than 15 days!!!
0
3.8K
232.1K
50.6K
Forward to community
DELTARUNE is 20% off for the Steam Winter Sale! We don't plan on discounting the game very often, so make sure to take this chance (to be a big shot)
0
355
29.8K
4.7K
Forward to community
Chapter 5 of DELTARUNE is now available as of June 24th!✨ Fans don't have to wait any longer for the newest installment of this game by the creator of UNDERTALE!🎮♥️ See #fanart# featuring #DELTARUNE# on #pixivision# today!👇
Show more
Toby Fox Has a First on the Billboard Charts With ‘Deltarune’ Soundtrack
Happy 2026 everybody! DELTARUNE Chapter 5 is on track to release later this year! The translation is already underway. Here are some lines I took from the translation sheet.
0
2.1K
133.8K
19.3K
Forward to community
It's Winter, so that means another DELTARUNE NEWSLETTER! This contains a lot of fun stuff, including some concept art for Chapter 3 and 4! We're sending it now, but it'll take a while for all the emails to go out.
Show more
0
679
69.4K
9.6K
Forward to community
🧩 Kimi K3’s MoE and Attention Are Built Around Trade-offs, Not Tricks Kimi K3’s open release has drawn attention to its scale. But its architecture tells a more useful story: the hardest part of scaling is keeping quality, efficiency, and stability in balance. Zhihu contributor 苏剑林 @Jianlin_S explains the design logic behind two core components: Stable LatentMoE and K3’s hybrid attention. At a high level: K3 = KDA + MLA + Stable LatentMoE + AttnRes 1️⃣ Stable LatentMoE: more experts at similar cost LatentMoE compresses each token into a smaller latent space before routing it to experts, then projects the result back to the full hidden dimension. This reduces expert computation and communication. The saved budget can support more, narrower experts without greatly increasing training or inference cost. But the longer projection chain also magnifies numerical instability. K3 introduces three fixes. 🔹 SiTU-GLU softly caps extreme activations in both branches of the expert network. Compared with hard clipping, soft capping preserves smoother optimization. 🔹 RMSNorm is placed before the final up-projection. It stabilizes training and helps balance routed experts against shared experts. 🔹 Quantile Balancing replaces the previous load-balancing update, which became unreliable as the expert pool grew. It approximates global routing quantiles with histograms, allowing efficient aggregation across machines. The broader lesson is clear: scaling MoE is not just about adding experts. Routing, activation ranges, normalization, and distributed communication must scale with them. 2️⃣ Why K3 still uses MLA Some newer models have moved away from MLA, partly because speculative decoding changes the inference trade-off. MLA keeps KV Cache small and remains highly competitive under fixed training and memory budgets. But its decoding path is relatively compute-heavy, leaving less room for Multi-Token Prediction to trade extra computation for speed. Other attention designs simply move the bottleneck: 🔹 Smaller designs may reduce computation but lose quality or require a larger KV Cache. 🔹 Larger designs can recover quality, but increase training and prefill costs. An ideal replacement would preserve quality, reduce KV Cache, lower decoding compute, and cost no more during training or prefill. No simple design currently satisfies all four conditions. K3 therefore keeps MLA and combines it with KDA. The linear-attention layers handle most long-context processing efficiently, while MLA preserves full-attention capacity where it matters. 3️⃣ “Abandoning MLA” is not so simple Architectures that appear to replace MLA may still retain its core intuition. For example, a wide MQA design with shared K and V resembles MLA’s decoding form. Sparsity and compression can then reduce its compute and cache costs. This can work, but it introduces more infrastructure complexity. So the current debate is less about whether MLA is obsolete. It is about which combination of full, linear, sparse, and compressed attention offers the best system-level trade-off. 4️⃣ Why K3 can remove RoPE K3 removes RoPE from its MLA layers. That would hurt a pure-MLA model. But K3 is a hybrid of KDA and MLA. KDA’s DeltaNet-style updates already introduce an implicit positional transformation. In this sense, KDA provides something similar to a generalized form of RoPE for the full network. Adding explicit RoPE back produced little difference, so K3 followed the simpler design. K3 is not truly position-free. Its positional structure is partly carried by KDA instead of an explicit embedding. ⚙ The real architecture lesson None of these choices is especially flashy in isolation. Stable LatentMoE controls the numerical and routing problems created by more experts. KDA and MLA divide long-context work according to their strengths. NoPE removes a redundant component only after the hybrid architecture makes it unnecessary. K3’s main design principle is therefore not novelty for its own sake. Every architectural change must justify itself across quality, efficiency, and stability. 🔗 Full reading: 📖Blog post: #KimiK3# #MoE# #Attention# #LLM# #AIInfra# #OpenSourceAI#
Show more