Register and share your invite link to earn from video plays and referrals.

Search results for OP_DROP_CGT
OP_DROP_CGT community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including OP_DROP_CGT
1001 Threads of Mizan is an action-adventure game with drop-in, drop-out co-op, letting you play solo or with up to 3 people. Developer MBC Game Studio told us how they developed the traversal and combat. Presented by 1001 Threads of Mizan
Show more
Qwen3.5-4B running CPU-ONLY at up to ~11+ tps on a Ryzen 5 laptop. No GPU. Running models in RAM and CPU is something the Qwen4 team is already thinking about. 💡 I built this standalone .exe because businesses are already deploying local models to save money and ensure privacy. (thumb drive friendly - just drag drop and doubleclick) My llama.cpp recipe 👉 🧠 Qwen3.5-4B Q4_K_M GGUF ⚙️ Ryzen 5 7540U — 6C/12T 🧵 --threads 9 🧵 --threads-batch 12 ⚡ --prio 2 🔄 --poll 50 📦 --batch-size 2048 📦 --ubatch-size 512 🚀 --flash-attn on 🧠 KV cache: q4_0 / q4_0 🔧 --repack 💾 --mmap 👤 --parallel 1 🚫 --device none 🚫 --gpu-layers 0 🚫 KV/op GPU offload 🚫 MTP OFF Interesting result 👉 MTP=3 was slower (~10 tps in benchmarks). Plain decode + 9 threads + Q4 KV hit ~11.4 tok/s. That's about 10% faster just from tuning llama.cpp — on a basic laptop CPU.
Show more
Question for the LLM Research Community: Is anyone aware of fully reproducible experimental results showing that MLA beats GQA under the same KV cache? I ask for two reasons: (1) I find it surprising that there is still a divide between Chinese labs using MLA, and Western labs using GQA + sliding window, when certainly many ablations have been run by many parties. (2) At Marin we are considering MLA for our next large scale run, but in early (!) smaller scale ablations it appears worse than heavily tuned feature-rich GQA, even after controlling for KV cache. I would like to run better experiments here. Below I cover my thoughts on general reproducibility, then specifics on MLA. Every empirical result in ML is only contextually true. Conditioned on the data distribution, optimizer settings, model width, model depth, finer architecture details, hardware, kernel engineering, initialization, token count, tokenizer, context length, and evaluation protocol, one can reach different conclusions. Contextual results are still useful. Typically if I see a promising method, I will first attempt a full 'context jump', where I apply it to my own context, hoping results transfer. Sometimes they do. If they don't, I can try 2 things: modify the implementation of the method, or modify the context. Ideally I have access to the full context of the original result. Then I can perform a 'context bridge', where I ablate one aspect of the context at a time, isolating exactly why a method performs differently. This lets me make an informed decision to either update my context to let the method shine, or stick with my context and leave the method out. MLA is tricky to assess at small scale. A core aspect of MLA is compressing hidden_dim->latent_dim. Then for each head, latent_dim->head_dim. Typically head_dim is fixed at 128, partially for hardware reasons, and partially for learning dynamics (head_dim of 8 wouldn't have sufficient representational capacity). To get MLA dynamics, you want hidden_dim>>latent_dim, and latent_dim>head_dim. This window closes at small scale. The degree of tuning can unfairly alter the scales. In GQA we have partial RoPE, QK Norm, Gated Attention, attention sharpening, sliding window, and other techniques that give a 30%+ training boost. They don't seem to give the same boost to MLA. On one hand, you want to compare techniques apples:apples with equal tuning. On the other hand, there is a finite amount of future tuning you can do, so prior tuning influences which approach is most pragmatic. Creating controlled tests between MLA and GQA is tricky. Several factors: kv_cache, quadratic attention flops, attention projection flops. kv_cache is controlled by scaling down kv_heads to match MLA, or scaling up kv_latent to match kv_heads. quadratic attention flops are controlled by scaling up GQA's query head count to match MLA head count, or scaling down MLA head count. Also scaling up GQA head_dim 128->192, or scaling down MLA head_dim to 192->128. In general, it's informative to context match to both option A's preferred context and option B's preferred context. Sliding window is another confounder. MLA is theoretically elegant, if we ignore RoPE. It replaces the 'replicate' op of kv_heads in GQA with a 'mix' op (pic below). Since the 'mix' can learn to 'replicate' if it wants, MLA is purely more expressive, and the cost of 'mix' is hidden at inference with absorb trick. Yet in practice, I find that at small scale this 'mix' op doesn't add much value and interacts poorly with the optimizer dynamics. And the change to RoPE hurts. My current plan is to first tune and ablate our model features around MLA, then run 3 scaling ladders: MLA, GQA with 2 kv_heads, and GQA with higher kv_heads. For each ladder, fit a loss vs compute projection. If MLA performs worse at our target compute compared to both GQA options, drop it. If MLA beats 2 kv_heads but loses to higher KV_heads, then it becomes a kv_cache tradeoff. Early results indicate MLA will perform worse than both feature-rich GQA ladders, but we will see. Any positive external reproducible results for MLA would help make sure I give it the best chance possible.
Show more
Op-ed: Salesforce just revealed the next battleground in AI — and it's not the models
OP Mainnet is running 3x the monthly transactions it did in early 2024. More on what's behind the volume later this week. Source: @Blockworks
OP-ED: A photograph taken at an elite conference in Sun Valley is going viral this week, and it may be the most perfect portrait of America’s new ruling class ever captured. There was Bari Weiss, the newly installed queen of safe heterodoxy — in a lab coat that screams “Women in STEM” — strolling through billionaire summer camp alongside Sheryl Sandberg, the patron saint of corporate feminism. Tucked under Sandberg’s arm was a copy of Democratic Pennsylvania Sen. John Fetterman’s book, Unfettered. Uniparty, anyone? The new uniparty is especially slippery because it claims to tolerate political disagreement — but only when that disagreement does not threaten the money, power or access of its members. Its members spar over pronouns, whether Trump goes too far when he attacks the media, or whether the new ballroom is a good idea. They declare themselves politically homeless and pretend to be the anti-establishment. But at the end of the day, they gather at the same invitation-only conferences, occupy the same boardrooms, circle through the same media outlets, and vote in the same direction. The Sun Valley photograph is so perfect because it captures the establishment’s response to populism. It didn’t retreat in horror. It tamed and re-packaged the revolt. Everyone gets to feel rebellious. No one important has to give anything up. Read more from @ambermarieduke below ⬇️
Show more
Op-Ed from @SecRubio: Why We’re Dismantling the ICC
0
621
4.2K
1K
Forward to community