Register and share your invite link to earn from video plays and referrals.

Rihard Jarc
@RihardJarc
Researching and investing in tech (AI, cloud, semiconductors, platforms). DMs open. Tweets are only opinions.
2.8K Following    81.3K Followers
A few thoughts on the $META AI model's progress, because I think it is significant. 1. It does seem that $META has now leapfrogged $GOOGL in model quality when it comes to Muse 1.2 for many use cases, which is very surprising given the timeframe. 2. This is still the "Muse Spark" family of models; $META already told us that bigger and more capable models (Watermelon) are coming. Given Muse Spark 1.2's performance already, the Watermelon family should be in the Fable category, which is impressive. 3. It does seem like $META is finally on a product scaling velocity curve (so the scaling foundations of their lab are set after 1 year of overhaul). And the shipping velocity is very good (3 releases in 4 months). 4. Given the recent rumored price hike of DeepSeek (if you use DeepSeek directly), it is clear that having enough compute to serve customers is critical. It doesn't help you if you have a great model, but most can't use it. $META is one of the few companies with compute capabilities comparable to, if not larger than, those of Anthropic or OpenAI.
Show more
Interesting with $META, based on the already guided CapEx for this year they are probably going to end 2027 with 7-10GWs of AI data center capacity. They could rent out 2-3GWs of that capacity, that would boost their net income by around $15-$30B (+20-40% of net income).
Show more
The signs were always there and as transparent as they could be... $META
The market reaction of punishing $META because of the increase in CapEx will turn out to be very wrong IMO. In the $META Q325 earnings call, Zuck already explained well that if $META overbuilds their compute infrastructure for internal needs, they can sell it to external parties. We are in one of the biggest compute demand/supply imbalances that seems to be getting worse, not better. $GOOGL GCP CEO Kurian expects demand to be bigger than supply for 10 years. $META's enormous AI compute capabilities will either return ROI in the form of their products (which they are already showing) or/and give great ROI by selling compute to external companies. $META investing in AI data center compute is a big asset, not a cash burn, similar to what metaverse investments were in large. The market is viewing it as a negative, instead of an enormous strategic advantage in a compute-constrained world.
Show more
Boom. $META is developing a cloud service to market excess AI processing capacity - Bloomberg
A new inference optimized hardware. They won’t have enough $TSM allocation to make serious market share gains in the near term, but this could shake the semis complex a lot. Strong backers. Looks like a complex build. Can help with the HBM bottleneck.
Show more
Introducing Cluster-Scale Memory (CSM) for low latency workloads. Today's AI chips using HBM can’t achieve SRAM-level decode speeds due to memory subsystem and interconnect bottlenecks. SRAM-only chips have lower FLOPs density and memory capacity, sacrificing throughput. You’re forced to make a tradeoff: serve at much slower speeds, or run at low batch sizes and suffer from higher costs. When running large MoE models, token routing across experts requires sending data through a deep memory hierarchy and a networking switch to reach a destination expert. Each memory layer inherently adds latency; thus, the best layer is no layer. We’ve designed a new architecture that creates a shared low-latency memory pool across the entire scale-up domain. We use a proprietary ultra-low-latency, high-bandwidth interconnect to enable dramatically faster memory access across chips. Our HBM/SRAM hybrid design solves both memory capacity and mem2mem latency, enabling high throughput and interactivity simultaneously. CSM improves latency and avoids today's cost, reliability, yield, thermal, and compute tradeoffs of SRAM-only chips, 3D DRAM chips, or optics.
Show more
An important read on the hyperscalers and why we are bullish.
0
94
4.2K
333
Forward to community
I just published my article on how hyperscalers like $AMZN, $MSFT, and $GOOGL will benefit in the token optimization era and why I believe the companies will re-rate much higher.
Show more
0
86
2.2K
263
Forward to community
Hyperscalers $AMZN $GOOGL and $MSFT are about to be the main beneficiaries of the token optimization trend. They get to squeez more revenue (tokens) from their existing infrastructure + they become the default orchestration layer for managing models, fine-tuning etc.
Show more
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
Show more
$MSFT, excluding its OpenAI stake, is trading at 17x forward P/E. Investors are betting heavily that the CEOs of hyperscalers don't know what they are doing.
Most investors still underestimate the value of chat surfaces when it comes to AI agents. In work agents Slack is the dominant channel and in consumer agents $META WhatsApp has a great position.
Introducing Claude Tag, a new way for teams to work with Claude. In Slack, Claude joins as a team member with access to the channels and tools you choose. Tag Claude in and delegate tasks to it while you focus on other work.
Show more
$META is announcing a new set of smartglasses at $299. It won't have a screen, but will include a camera and personal speakers with Meta AI. A great launch moment and a very accessible price point for people to try. If the AI is good enough, people will get hooked on it.
Show more
Reminder that the market can be extremely irrational even in the big cap names. I view today’s period as similar.
$META is now trading at 11x EV/FCF (TTM). Can't wait to look back at this in 2 years.
Didn't think I would see $MSFT trading at a 20x forward P/E again in the $MSFT "Azure era" now coming at the April war lows. The interesting part is that growth for Azure looks as good as it has ever been (+40%) with a massive backlog and severe global compute shortage. On the software part, there are no real signs of companies canceling Office licenses (despite the fears around it).
Show more
A MUST-read interview with a Siemens employee explaining just how high demand is for energy equipment right now because of AI: 1. The whole situation is shocking even for people who have been in the business for 40 years. They are getting orders that are double the size of what their entire factory can produce in a year. 2. Demand is so high in the last 5-8 months that they don't need to convince or send any analysis (such as CO2 emissions, etc.) to clients because they just want the equipment, because there's so much backlog that they just want to catch the order. 3. Decisions are being made very quickly by clients; the backlog for some of the energy equipment companies is 5-6 years. For transformers, the situation is even more difficult. 4. He mentions that right now, data center builders do not care about sustainability; they just want power at any expense, reliable power. They say they will think about sustainability later. 5. The orders have gone from previous 20-30 MW orders to now 200-500 MW units. Customers have previously wanted to get equipment from different OEMs, but now they prefer an integrated standardized solution. 6. An interesting dynamic is that even though the data center requires 100 MW, the builders are buying N+1 units of gas turbines (so more than just for 100 MW) as backups, as well as having more energy capacity, as they believe they will continue to grow that data center. 7. He does believe there is some double booking going on on transformers and switchgears because of extra-long lead times. 8. Everyone is trying to reduce PUE, and water use effectiveness, but even after improving, they just use the same power to run more compute. 9. The problem is also liquid cooling, as it is expensive, and water availability in many regions is a problem. 10. Margins on equipment in the sector have gone from 4-6%, where they were 2-3 years ago, to 20-23% and in some cases even 40%. The data center builders know the margins are high, but they are fine with it because they just want to get it. found on @AlphaSenseInc
Show more
The market reaction of punishing $META because of the increase in CapEx will turn out to be very wrong IMO. In the $META Q325 earnings call, Zuck already explained well that if $META overbuilds their compute infrastructure for internal needs, they can sell it to external parties. We are in one of the biggest compute demand/supply imbalances that seems to be getting worse, not better. $GOOGL GCP CEO Kurian expects demand to be bigger than supply for 10 years. $META's enormous AI compute capabilities will either return ROI in the form of their products (which they are already showing) or/and give great ROI by selling compute to external companies. $META investing in AI data center compute is a big asset, not a cash burn, similar to what metaverse investments were in large. The market is viewing it as a negative, instead of an enormous strategic advantage in a compute-constrained world.
Show more
$MU almost 75% gross margin and a 81% gross margin guide for next Q tells you everything you need to know about who has the power in this AI hardware super cycle.
$META is now trading at 11x EV/FCF (TTM). Can't wait to look back at this in 2 years.