Register and share your invite link to earn from video plays and referrals.

Search results for smooth_ad
smooth_ad community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including smooth_ad
I added to $AEHR today as well. Here's my thesis why: The growth inflection: FY2027 revenue guided to $130–150M vs. FY2026's $50M — 160–200% growth Guidance carries an 18–22% pre-tax margin, so this is profitable growth, not revenue bought with margin Demand is contracted, not hoped for: Record Q4 bookings of $60.7M, up ~500% YoY Effective backlog of $100.6M — covering ~72% of the FY27 guidance midpoint The end-market pivot is the real story ~95% of revenue now from AI processors, silicon photonics and power semis — versus 95% EV silicon carbide two years ago AI expected to be ~70% of FY2027 revenue, silicon photonics ~15% Valuation is the main pushback: ~21x guided FY27 sales at recent prices, and one bear case puts it at 59x forward FY2028 profit. I expect lumpy prints, not a smooth ramp but firmly believe this will be a large company in 2028! NFA. DYOR
Show more
.@benedictpolizzi's compliments AND skin are smooth on the Grammys red carpet 💫 (ad)
0
41
1.7K
60
Forward to community
This isn't The Addams Family! Still controllable even after leaving the body! China-made intelligent bionic hand makes its debut at WAIC2026, with lifelike synthetic skin that looks almost real, and "remote" control that is smooth and seamless. Netizens: "Detachable hand" becomes a reality — Cool! #Technology# #IntelligentBionics# #MadeInChina# #AI#
Show more
Fourier showed that any periodic function can be built by adding sine waves. For example take a square wave: it follows a repeating up-down pattern over a fixed period. Even though sine waves are smooth, we can stretch, scale, and shift them. By adding multiple sine waves with different properties, we start to approximate the square wave. At first, it’s rough—but as we include more terms, the sum becomes increasingly accurate, eventually closely matching the square wave.
Show more
📈 Bigger Recommender Transformers Do Not Automatically Scale Meta’s recent ads-ranking work has brought recommender scaling back into focus. Its central lesson is simple: predictable gains appear only when the model, user sequences, features, and serving system scale together. Zhihu contributor 九老师 explains why Transformer scaling is fundamentally a scale-matching problem, not a race to add more parameters. 1️⃣ Scaling changes the entire training system Increasing depth or width changes far more than model capacity. It also changes residual magnitudes, gradient flow, attention logits, optimizer states, and the best learning rate. A larger model can therefore appear to train normally while some layers contribute very little. The real question is not “How large is the Transformer?” but “Can every part of the system remain effective at this scale?” 2️⃣ Normalization determines whether depth is useful Post-Norm Transformers become increasingly fragile as depth grows. Gradient spikes can distort optimizer states, attention can collapse, and lower layers may gradually stop learning. Pre-Norm creates a cleaner gradient path and is generally more stable. But it has its own limitation: as residual signals accumulate, later layers may become dominated by the identity path. The network becomes deeper without learning proportionally richer representations. This is why residual scaling and initialization must evolve with model depth. 3️⃣ Attention and optimization must scale too Attention can fail in two opposite ways: 🔹 Entropy collapse: attention becomes extremely sharp and concentrates on very few positions. 🔹 Rank collapse: repeated mixing makes token representations increasingly similar. Techniques such as QK normalization can control attention sharpness, while residual scaling helps preserve token-specific information across layers. Optimizer settings also cannot be copied blindly from smaller models or LLM recipes. Recommendation data changes quickly, so historical gradients may become stale faster. Model scale, learning rate, initialization, and optimizer dynamics must be tuned as one system. 4️⃣ Loss alone cannot reveal silent failures A smooth training curve does not prove that the full Transformer is learning. Useful internal signals include: 🔹 Gradient strength across different layers 🔹 The size of parameter updates relative to parameter weights 🔹 Similarity between adjacent-layer representations 🔹 Attention-logit magnitude and attention entropy These diagnostics reveal whether lower layers are inactive, attention is collapsing, or updates have become poorly scaled. 5️⃣ Bigger models need richer inputs Scaling parameters alone may bring little improvement when the model still receives heavily compressed or low-complexity features. The author recalls that multimodal embeddings initially produced little gain in one recommender system. A manually designed distance feature worked better. After the model gained enough capacity, those raw embeddings became much more useful. The lesson is not that larger Transformers always win. It is that model capacity and input complexity must grow together. If the available information is simple, a smaller architecture may still be the better choice. ⚙ The core lesson A recommender Transformer is not a plug-and-play module. Features, sequences, tokenization, normalization, attention, optimization, training, and serving all interact. Scaling one component while freezing the others often creates cost without real capability. Using a Transformer is not the same as building a Transformer-native recommender system. A model truly scales only when the whole system scales with it. 🔗 Recent context: 🔗 Full Reading: #RecommenderSystems# #Transformers# #ScalingLaws# #MachineLearning# #AIInfra#
Show more
As uncle @JensenHuang would say, that loser premise makes no sense to me. Capability doesn't advance only as a smooth race for the biggest general model or the most FLOPs. The frontier is jagged, and the cleanest illustration is the strawberry problem. Ask a model that can hold its own with a physics PhD how many R's are in the word "strawberry" and for a long stretch it would answer two, confidently, every time. The same system could work through a graduate-level proof and then insist 9.11 is a bigger number than 9.9. Each specific embarrassment gets patched as it goes viral. The shape underneath survives every patch, that is - enormous strength in some places, sudden collapse in others, and no way to tell which ground you're standing on until you test it. But counting letters and numbers are funny failures. The expensive version of the same problem shows up in security. Hugging Face was mid-incident, working an agentic attack, and needed a model to review logs. The commercial frontier models refused due to "safety" reasons. From inside a guardrail, a defender reading attack traffic and an attacker reading attack traffic look identical, so the safest available behaviour is to decline. The team switched to an open-weights model on their own infrastructure and got the analysis done. The most capable system in the room was the one that couldn't help, at the exact hour it was needed. That refusal is one of many permanent structural feature of general-purpose models, and it sits right on top of a business. Aether AI operates in the gap between those two failures, and @tryaether_ai AI is Australian. So telling Australian founders the game is over at the model layer, dig up some lithium and prepare a robot tax, is exactly the wrong message. The scoreboard was never general-model scale. It's about who owns the jagged edges that customers, regulators, and national security actually pay for. Australia can own plenty of them. Some of us already do.
Show more
No one expects a polished elevator pitch at a conference, but showing up prepared for smooth, natural conversations? That’s a massive advantage. Great reminder by our Community Architect @Maky_sol. Learn with the Empire. ⚔️
Show more