Register and share your invite link to earn from video plays and referrals.

Search results for ScalingLaws
ScalingLaws community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including ScalingLaws
📈 Bigger Recommender Transformers Do Not Automatically Scale Meta’s recent ads-ranking work has brought recommender scaling back into focus. Its central lesson is simple: predictable gains appear only when the model, user sequences, features, and serving system scale together. Zhihu contributor 九老师 explains why Transformer scaling is fundamentally a scale-matching problem, not a race to add more parameters. 1️⃣ Scaling changes the entire training system Increasing depth or width changes far more than model capacity. It also changes residual magnitudes, gradient flow, attention logits, optimizer states, and the best learning rate. A larger model can therefore appear to train normally while some layers contribute very little. The real question is not “How large is the Transformer?” but “Can every part of the system remain effective at this scale?” 2️⃣ Normalization determines whether depth is useful Post-Norm Transformers become increasingly fragile as depth grows. Gradient spikes can distort optimizer states, attention can collapse, and lower layers may gradually stop learning. Pre-Norm creates a cleaner gradient path and is generally more stable. But it has its own limitation: as residual signals accumulate, later layers may become dominated by the identity path. The network becomes deeper without learning proportionally richer representations. This is why residual scaling and initialization must evolve with model depth. 3️⃣ Attention and optimization must scale too Attention can fail in two opposite ways: 🔹 Entropy collapse: attention becomes extremely sharp and concentrates on very few positions. 🔹 Rank collapse: repeated mixing makes token representations increasingly similar. Techniques such as QK normalization can control attention sharpness, while residual scaling helps preserve token-specific information across layers. Optimizer settings also cannot be copied blindly from smaller models or LLM recipes. Recommendation data changes quickly, so historical gradients may become stale faster. Model scale, learning rate, initialization, and optimizer dynamics must be tuned as one system. 4️⃣ Loss alone cannot reveal silent failures A smooth training curve does not prove that the full Transformer is learning. Useful internal signals include: 🔹 Gradient strength across different layers 🔹 The size of parameter updates relative to parameter weights 🔹 Similarity between adjacent-layer representations 🔹 Attention-logit magnitude and attention entropy These diagnostics reveal whether lower layers are inactive, attention is collapsing, or updates have become poorly scaled. 5️⃣ Bigger models need richer inputs Scaling parameters alone may bring little improvement when the model still receives heavily compressed or low-complexity features. The author recalls that multimodal embeddings initially produced little gain in one recommender system. A manually designed distance feature worked better. After the model gained enough capacity, those raw embeddings became much more useful. The lesson is not that larger Transformers always win. It is that model capacity and input complexity must grow together. If the available information is simple, a smaller architecture may still be the better choice. ⚙ The core lesson A recommender Transformer is not a plug-and-play module. Features, sequences, tokenization, normalization, attention, optimization, training, and serving all interact. Scaling one component while freezing the others often creates cost without real capability. Using a Transformer is not the same as building a Transformer-native recommender system. A model truly scales only when the whole system scales with it. 🔗 Recent context: 🔗 Full Reading: #RecommenderSystems# #Transformers# #ScalingLaws# #MachineLearning# #AIInfra#
Show more
An AI agent's performance is governed not by how much it computes, but by how well that compute turns into good feedback 📈 Title: Scaling Laws for Agent Harnesses via Effective Feedback Compute URL: 📈 Overview This work proposes Effective Feedback Compute (EFC), a metric that reframes agent scaling efficiency around feedback quality rather than raw compute. It measures whether computation actually improved decisions. ❓ Challenges Solved We tend to reason about performance via raw metrics — tokens, tool calls, cost. But these mask whether feedback truly improved decision-making. Redundant, invalid, or unused feedback doesn't help. 💡 Methodology & Proposed Approach ・EFC credits feedback only when it is informative, valid, non-redundant, and retained for later decisions ・It normalizes by task demands for fair cross-task comparison ・Evaluated on synthetic tasks, code tasks, real traces, and prospective tests, vs raw-compute and SAS baselines 📊 Experimental Results EFC's explanatory power stood out (R² vs performance). ・Raw tokens/tool calls: R²=0.33-0.42 ・SAS baseline: 0.88, Oracle-EFC: 0.94, task-normalized: 0.99 ・Real traces: 0.92, prospective holdout: 0.85 ・Matched-budget interventions that improved feedback quality lifted success from 0.27 to 0.90 #AIAgents# #ScalingLaws#
Show more
Abra: Scaling Diffusion Image Training A comprehensive scaling laws study of text-to-image diffusion models from Luma AI: "We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute (1019 to 1022 FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately 200 image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs." paper link:
Show more
Deconstructing Scaling Laws: The Triad of Optimization, Architecture, and Data
The bear case (other than scaling laws failing) for Ant + OAI was previously that Google would weaponize gross margins and commoditizes the frontier. The bear case going forward is that republicans stay firmly in control past 2028 and Elon leverages a compute advantage to spin faster cycles on his RSI and unlike the Google org has the risk tolerance to make broadly available a frontier model that the labs are too scared to. We are approaching the point where risk tolerance is very very important. Suspect “safety & alignment” is going to take up a much larger portion of opex for the labs vs what people might think. Think of the five worst people you have ever met. Now imagine if you gave each of them ten thousand employees and free rein. Terrifying. (Example stolen from latest moonshots pod)
Show more
At Marin we’ve found that scaling laws let us simulate the entire training trajectory. This became clear in our last 67B MoE run. The run hit the loss target within 1%, but perhaps more interestingly, arbitrary points during training could also be predicted to within 1%. This model is finishing up long context and will soon enter RL. Based on this finding, we’ve used a scaling ladder to simulate the training trajectory of the 535B MoE we kicked off yesterday. Will see how it goes. Now if someone can figure out how to fit a scaling law to each parameter we can just predict our final model and skip this training business. Something else I've learned from this process is that getting a training run off the ground is way more than just fiddling with the architecture. Lot of herculean efforts from teammates to get trillions of tokens curated, get hardware running on a new GPU stack, improve MFU, and propose great ideas that were tested and integrated. Given that this is our largest run to date, I anticipate lots of learning and adapting along the way. But thats part of the fun. Details:
Show more
capital efficiency unlocks some very interesting scaling laws
NVIDIA CEO Jensen Huang says one scaling law multiplies AI faster than NVIDIA can hire engineers. Most people know three AI scaling laws. Pre-training. Post-training. Test-time. Each one multiplies intelligence by throwing more compute at a different stage. Jensen Huang says there's a fourth and it's the one that will dominate... Agentic scaling law. "During test time, that agentic system goes off and does research, bangs on databases, uses tools," Huang says. "And one of the most important things it does is spawn off a whole bunch of sub-agents." That's the multiplier. One AI worker can become a team. Then a department. Then a company. "It's so much easier to scale NVIDIA by hiring more employees than it is to scale myself," Huang says. Now imagine scaling without a payroll constraint. "The agentic scaling law — it's kind of like multiplying AI," Huang says. "We could spin off agents as fast as you want to spin off agents." Each agent spins off sub-agents. Each sub-agent spins off more. The compute requirement compounds inside a single query. And every agent generates new data, new experiences, new edge cases. "Wow, this is really good. We ought to memorize this," Huang says. "That data set comes back to pre-training." The four scaling laws don't compete. They feed each other. Agentic systems produce data, which feeds pre-training, which smartens the base model, which enables better agents, which produce more data. A flywheel that compounds forever. The companies pricing in three scaling laws are mispricing the fourth. The fourth eats the other three for lunch. P.S. Pull the thread on any story like this and you'll find the hidden incentive at the other end. As Munger said: "Show me the incentive and I'll show you the outcome." So I wrote a short book on how to spot them and design your own. Comment "INCENTIVES" and I'll send you the details. If you're new here, follow @GeniusGTX for content on the greatest minds in economics, psychology, and history. — Jensen Huang ( @nvidia ), NVIDIA CEO, on Lex Fridman's ( @lexfridman ) podcast
Show more
everyone thinking about agi, scaling laws, and the implications of superintelligence should be watching pantheon right now
not for anything but it’s structurally incoherent to believe scaling laws hold without also believing we are about to enter a period where infrastructure attacks on hyperscale compute and its factor inputs (power, chips) are incredibly common distillation is not the frontier
Show more