Register and share your invite link to earn from video plays and referrals.

Search results for QK新着動画
QK新着動画 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including QK新着動画
vLLM tops the Artificial Analysis leaderboard 🎉 vLLM tops @ArtificialAnlys on DeepSeek V3.2 and ranks among the top deployments of MiniMax-M2.5 and Qwen 3.5 397B. The leading deployments of these models are now open source. How each result was built: 🔹 DeepSeek V3.2 — Aggressive op fusion across the attention path collapsed ~33 per-layer kernels down toward ~10. 🔹 MiniMax-M2.5 — Custom EAGLE3 draft trained against the target's own token distribution via TorchSpec, plus a custom QK-norm fusion for MiniMax's TP-aware attention. 🔹 Qwen 3.5 397B — Targeted fusions plus a QK-norm fix for Qwen's linear-attention path. Every optimization is in vLLM main or on its way upstream. Huge thank you to @inferact, @digitalocean, @nvidia, @RedHat_AI, and the vLLM community 🙏 Full breakdown 👇
Show more
📈 Bigger Recommender Transformers Do Not Automatically Scale Meta’s recent ads-ranking work has brought recommender scaling back into focus. Its central lesson is simple: predictable gains appear only when the model, user sequences, features, and serving system scale together. Zhihu contributor 九老师 explains why Transformer scaling is fundamentally a scale-matching problem, not a race to add more parameters. 1️⃣ Scaling changes the entire training system Increasing depth or width changes far more than model capacity. It also changes residual magnitudes, gradient flow, attention logits, optimizer states, and the best learning rate. A larger model can therefore appear to train normally while some layers contribute very little. The real question is not “How large is the Transformer?” but “Can every part of the system remain effective at this scale?” 2️⃣ Normalization determines whether depth is useful Post-Norm Transformers become increasingly fragile as depth grows. Gradient spikes can distort optimizer states, attention can collapse, and lower layers may gradually stop learning. Pre-Norm creates a cleaner gradient path and is generally more stable. But it has its own limitation: as residual signals accumulate, later layers may become dominated by the identity path. The network becomes deeper without learning proportionally richer representations. This is why residual scaling and initialization must evolve with model depth. 3️⃣ Attention and optimization must scale too Attention can fail in two opposite ways: 🔹 Entropy collapse: attention becomes extremely sharp and concentrates on very few positions. 🔹 Rank collapse: repeated mixing makes token representations increasingly similar. Techniques such as QK normalization can control attention sharpness, while residual scaling helps preserve token-specific information across layers. Optimizer settings also cannot be copied blindly from smaller models or LLM recipes. Recommendation data changes quickly, so historical gradients may become stale faster. Model scale, learning rate, initialization, and optimizer dynamics must be tuned as one system. 4️⃣ Loss alone cannot reveal silent failures A smooth training curve does not prove that the full Transformer is learning. Useful internal signals include: 🔹 Gradient strength across different layers 🔹 The size of parameter updates relative to parameter weights 🔹 Similarity between adjacent-layer representations 🔹 Attention-logit magnitude and attention entropy These diagnostics reveal whether lower layers are inactive, attention is collapsing, or updates have become poorly scaled. 5️⃣ Bigger models need richer inputs Scaling parameters alone may bring little improvement when the model still receives heavily compressed or low-complexity features. The author recalls that multimodal embeddings initially produced little gain in one recommender system. A manually designed distance feature worked better. After the model gained enough capacity, those raw embeddings became much more useful. The lesson is not that larger Transformers always win. It is that model capacity and input complexity must grow together. If the available information is simple, a smaller architecture may still be the better choice. ⚙ The core lesson A recommender Transformer is not a plug-and-play module. Features, sequences, tokenization, normalization, attention, optimization, training, and serving all interact. Scaling one component while freezing the others often creates cost without real capability. Using a Transformer is not the same as building a Transformer-native recommender system. A model truly scales only when the whole system scales with it. 🔗 Recent context: 🔗 Full Reading: #RecommenderSystems# #Transformers# #ScalingLaws# #MachineLearning# #AIInfra#
Show more
Question for the LLM Research Community: Is anyone aware of fully reproducible experimental results showing that MLA beats GQA under the same KV cache? I ask for two reasons: (1) I find it surprising that there is still a divide between Chinese labs using MLA, and Western labs using GQA + sliding window, when certainly many ablations have been run by many parties. (2) At Marin we are considering MLA for our next large scale run, but in early (!) smaller scale ablations it appears worse than heavily tuned feature-rich GQA, even after controlling for KV cache. I would like to run better experiments here. Below I cover my thoughts on general reproducibility, then specifics on MLA. Every empirical result in ML is only contextually true. Conditioned on the data distribution, optimizer settings, model width, model depth, finer architecture details, hardware, kernel engineering, initialization, token count, tokenizer, context length, and evaluation protocol, one can reach different conclusions. Contextual results are still useful. Typically if I see a promising method, I will first attempt a full 'context jump', where I apply it to my own context, hoping results transfer. Sometimes they do. If they don't, I can try 2 things: modify the implementation of the method, or modify the context. Ideally I have access to the full context of the original result. Then I can perform a 'context bridge', where I ablate one aspect of the context at a time, isolating exactly why a method performs differently. This lets me make an informed decision to either update my context to let the method shine, or stick with my context and leave the method out. MLA is tricky to assess at small scale. A core aspect of MLA is compressing hidden_dim->latent_dim. Then for each head, latent_dim->head_dim. Typically head_dim is fixed at 128, partially for hardware reasons, and partially for learning dynamics (head_dim of 8 wouldn't have sufficient representational capacity). To get MLA dynamics, you want hidden_dim>>latent_dim, and latent_dim>head_dim. This window closes at small scale. The degree of tuning can unfairly alter the scales. In GQA we have partial RoPE, QK Norm, Gated Attention, attention sharpening, sliding window, and other techniques that give a 30%+ training boost. They don't seem to give the same boost to MLA. On one hand, you want to compare techniques apples:apples with equal tuning. On the other hand, there is a finite amount of future tuning you can do, so prior tuning influences which approach is most pragmatic. Creating controlled tests between MLA and GQA is tricky. Several factors: kv_cache, quadratic attention flops, attention projection flops. kv_cache is controlled by scaling down kv_heads to match MLA, or scaling up kv_latent to match kv_heads. quadratic attention flops are controlled by scaling up GQA's query head count to match MLA head count, or scaling down MLA head count. Also scaling up GQA head_dim 128->192, or scaling down MLA head_dim to 192->128. In general, it's informative to context match to both option A's preferred context and option B's preferred context. Sliding window is another confounder. MLA is theoretically elegant, if we ignore RoPE. It replaces the 'replicate' op of kv_heads in GQA with a 'mix' op (pic below). Since the 'mix' can learn to 'replicate' if it wants, MLA is purely more expressive, and the cost of 'mix' is hidden at inference with absorb trick. Yet in practice, I find that at small scale this 'mix' op doesn't add much value and interacts poorly with the optimizer dynamics. And the change to RoPE hurts. My current plan is to first tune and ablate our model features around MLA, then run 3 scaling ladders: MLA, GQA with 2 kv_heads, and GQA with higher kv_heads. For each ladder, fit a loss vs compute projection. If MLA performs worse at our target compute compared to both GQA options, drop it. If MLA beats 2 kv_heads but loses to higher KV_heads, then it becomes a kv_cache tradeoff. Early results indicate MLA will perform worse than both feature-rich GQA ladders, but we will see. Any positive external reproducible results for MLA would help make sure I give it the best chance possible.
Show more
New NanoGPT Speedrun WR at 76.3 (-2.9s) from @Lisennlp , with a radical boost to the final attention layer that drops step count by 7%, called lightweight Dynamically Composable MHA. A secondary local window (112 tokens) is recomputed on the same q/k. After softmax it mixes the per-head attn maps into one unified map, then applies the shared map to each head's values with a per-head scale. Every mix and scale operation is computed per-token from a MUDD MLP. All in a custom kernel. Takes inspiration from Dynamically Composable MHA
Show more
[Attention] Attention is how one token pulls relevant information from itself and the tokens before it. At each layer, learned projections turn each token’s current representation into three vectors: a query, a key, and a value. The query represents what the current position is looking for. Keys describe what each token can be matched on, while values carry the information that can be pulled in. For “runs,” one attention head compares its query with the keys of “The,” “chip,” “Alice,” “designed,” and “runs” itself. Each comparison is a scaled dot product, q · k / √d. Softmax turns the scores into weights that add up to one, and the values are multiplied by those weights and summed. A transformer runs several heads in parallel. Each has its own learned projections, so it can view the same tokens differently. Their outputs are combined and passed through the remaining layers. A head may emphasize the subject while another captures a different relationship, but these are learned tendencies, not roles assigned in advance. But where did all those keys and values come from? And when the next token arrives, does the model have to build the earlier ones all over again? It does not. Tomorrow: KV cache. The paper that introduced the transformer: Vaswani et al., NeurIPS 2017.
Show more
Attention is a lookup. Each token builds a query, compares it against every key in the sequence, and pulls value vectors weighted by the match. Stack that 96 layers deep and you get a frontier model. Video covers the full pipeline: Q/K/V, attention scores, encoder blocks.
Show more