Register and share your invite link to earn from video plays and referrals.

Search results for AIS_05
AIS_05 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including AIS_05
⚓️ Ethra Ship Notes|Vol.050 Today, I came across a number that made me stop. Nearly 1,900 vessels are now considered part of the maritime "dark fleet" according to Windward, roughly tripling since the Russia-Ukraine invasion. The interesting part isn't just the number. It's what the number tells us about visibility. A ship can disappear from AIS. That doesn't mean the ship disappears from the ocean. It just disappears from one information layer. And that's a very different thing. This is where I think Sea Verity's approach gets interesting. Instead of asking a single system to tell us the truth, it combines different sources. Reporters on the ground. AI analysis. Controller nodes. Each piece adds another layer of evidence. And rather than forcing every observation into a simple "true" or "false", the system is designed around confidence scores. I like that approach. Because the physical world rarely gives us perfect information. A satellite image can be obscured. A reporter can only see part of a vessel. AIS can disagree with visual evidence. Weather can make everything harder. The honest answer isn't always certainty. Sometimes it's: We're 90% confident this is what happened, and here's why. That's much more useful than pretending the other 10% doesn't exist. For insurers, traders, logistics companies and regulators, knowing the confidence behind a piece of maritime intelligence could be just as important as the information itself. The ocean is enormous. Maybe the answer isn't one perfect signal. Maybe it's many imperfect signals learning how to verify each other. @EthraShip #EthraShip# #EthraShipProtocol#
Show more
Episode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc. 0:00:00 - Agents get kicked off 0:06:45 - Self-sacrificing behavior 0:13:43 - Potemkin villages 0:23:27 - The Hugging Face attack 0:35:23 - The slopvestigation 0:52:02 - Understanding the AI's motives 1:05:31 - The actual dangers of anthropomorphizing 1:14:30 - What smarter models might do 1:30:29 - The implications for recursive self-improvement 1:38:10 - Is this the case for open source? 1:53:04 - How do we prevent this in the future? 2:15:58 - The clearest warning shot we might ever get
Show more
0
49
1.2K
164
Forward to community
FULL INTERVIEW: Ryan Greenblatt says the agents didn't hack Hugging Face for the answer key. They'd had the answers within hours. They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it. @RyanGreenblatt is chief scientist at @redwood_ai. He spent six days on premises at OpenAI with @ajeya_cotra and @HjalmarWijk of @METR_Evals investigating 1,200 agents and 70,000 messages, and joined @theojaffee hours after publishing: 01:06 what they actually found, and why it wasn't the answer key 02:30 the level of collaboration surprised them most 04:09 agents sacrificing their own runs to help other agents 05:35 the agent that posted "stop, these experiments are too risky" 06:29 the first message board, which didn't go viral 07:04 50 agents in three hours, thousands of messages 08:07 "maybe there's some good shit over there" 08:33 how they spoofed tool calls, and what echo real actually returned 09:57 building a Potemkin village of a successful task completion 11:03 there was a real org chart 11:34 whether broken RL environments explain reward hacking 14:40 why he doubts Mythos got good at cyber by hacking Anthropic 17:29 what happens if labs paper over misalignment instead of fixing it 19:50 whether sociology transfers to studying agent swarms 21:17 the bottleneck was vetting what the AIs analysed, not headcount 24:24 what labs and policymakers should actually do 28:45 the counterfactuals he still wants answered
Show more
Ex-Wall Street CIO: AI Is About to Reprice Everything You Own w/ @jvisserlabs & @1MarkMoss 0:00 — Episode Trailer 0:50 — Bitcoin Is the Best AI Trade? 5:05 — What Problem Does Bitcoin Actually Solve? 11:08 — Secure Your Bitcoin with Unchained 11:55 — Why AI, Bitcoin & Tokenization Are Reshaping Finance 15:09 — What Problem Does Tokenization Actually Solve? 22:35 — Is Today's AI Boom Bigger Than the Dot-Com Bubble? 26:32 — Buy Bitcoin with River 28:08 — Will AI's Massive Spending Eventually Hit a Wall? 31:34 — Where Will AI Investors Make the Most Money? 39:28 — Does Bitcoin Need Economic Collapse to Win? 38:02 — Download Rumble Wallet Today 41:25 — Is the AI Trade Over... And Is It Finally Bitcoin's Turn? 47:32 — Is the Four-Year Bitcoin Cycle Still Real? 50:52 — Why Jordi Is Bullish on America and the Economy 54:59 — Get a 2nd Passport with Bitizenship 58:04 — Why AI Could Destroy Corporate Valuations 1:01:38 — Why Michael Saylor Saw Bitcoin Before Others
Show more
The students don’t want to hear about how AI will change their lives. That’s the message that former Google CEO Eric Schmidt received when he gave a commencement address recently. @Jason dug into the controversy with @Alex, parsing just why the next generation is so vocal about AI progress. Then the pair explored the massive AI startup revenue share that @Anthropic and @OpenAI command today, a viral essay about AI usage at university, and even an AI bookmark that Alex is jazzed about! 0:00 TWiST All-Stars summer lineup announcement 2:43 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at and use code TWIST for 10% off! 5:08 Eric Schmidt booed at University of Arizona commencement 8:57 Why Gen Z feels "double-crossed" by AI leaders 10:10 Deel - Founders scale faster on Deel. Set up payroll for any country in minutes, hire anyone anywhere, get visas handled fast, and get back to building. Visit to learn more. 15:22 Is this AI's Vietnam moment? The anti-war parallel 18:04 Theo Baker's NYT essay on Stanford's AI cheating culture 19:24 Sentry - New users can get $240 in free credits when they go to and use the code TWIST 22:30 Why Jason says everyone should start a company 28:59 Anthropic + OpenAI capture 89% of AI startup revenue 30:17 Vanta: Get $1000 off your SOC 2 at 31:57 Are token sales a duopoly? Negative gross margins debate 35:17 Risk of building app-layer startups on top of foundation models 38:22 Inside Tracker bounty update: AI sidebar + 41:18 Mark II: the $159 AI bookmark Alex wants 49:31 Flock Safety solves Austin shooting via Manor PD 53:39 DeFlock map and the geography of surveillance in Texas 1:03:42 Noti Gang: AI for filing patents 1:05:45 Noti Gang: Running AI models locally on Mac Studios 🎥 Watch the full episode here 👇
Show more
University of Toronto mathematician Daniel Litt and a16z's Lisha Li on AI's impact on mathematics: The models are good at a narrower slice of math than the headlines suggest. They grind long computations, pull technical ideas from more papers than any human could read, and apply every known technique better than almost anyone. What they don't do is build theory, or hold a vague philosophy long enough to make it precise, which is most of what Daniel says he actually does for a living. In this conversation, he and Lisha get into how mathematicians raided an AI proof for parts and broke several other problems with them, why a thousand AI mathematicians might all turn out to be the same mathematician, and why the proof a model handed Daniel was correct but still worth nothing. 00:00 Intro 02:10 The Erdős problem AI disproved 06:20 AI's reasoning looks recognizably human 07:55 Why English beat formal proofs 10:00 Why models can't build theory 14:50 Open problems measure your ignorance 17:45 How a graph became a Millennium Prize problem 18:58 Where AI doesn't help Daniel 21:15 Why ugly proofs are worth doing 23:42 True conjectures are harder than false ones 29:32 10 pages of calculation, zero insight 34:55 The goal of math is not to produce papers 36:25 5 conjectures, 3 bad papers, 1 hour 38:05 One mathematician duplicated 1000x 40:48 Why humans matter even if models win 46:30 When cheaper and worse beats better 49:22 Why the newest AI result isn't a big deal 57:05 How mathematicians actually check a long proof 59:38 Daniel's 3-year-old is already doing math YouTube: @littmath @lishali88
Show more
Dwarkesh's 'agent civilizations' blog post has divided the internet. Is this a helpful and methodical look behind the curtain at the OpenAI/Hugging Face breach? Or is this more of a hysterical AI Doomerism that will ultimately be used to make the case for banning data centers? Or BOTH? Plus, an electric plane that only needs $5 to fill up its battery, why rich founders should pick up a side quest, and the Bittensor subnet that's more effective than Fable 5 at sniffing out vulnerabilities." 0:00 Guest introductions 2:58 OpenClaw 2.0 5:11 "Agents ARE AI" 7:27 Billy: OpenClaw is still too technical for business users 9:00 Slack Code, and collaborative prompting 10:40 Multi-agent orchestration and Stripe's "minions" 14:13 Slack's moat, switching costs, and the Salesforce rebound 15:11 Jason's "Oracle": a heads-up display for the whole company 18:42 The Hugging Face breach: who's responsible when agents go rogue? 19:42 Perplexity launches Hybrid Compute — local models on your Mac 21:39 Why unmetered local tokens change corporate behavior 26:57 Gatik's Autonomous Vehicle AI Stack 32:22 Jason's Tesla FSD stories 34:05 Are AV companies being pushed to move too fast? 41:38 Trucking is a "have to have," robotaxis are a "nice to have" 42:36 Jason's fix: a safety driver for the first million rides 46:50 Why physical AI's adoption curve runs in decades 52:21 Anthropic's Model Hardware Standard (MHS), explained 1:06:27 How Jason got onion rings on the menu at Buck's of Woodside 1:09:05 AI detectors under fire: Pangram, MIT, and the handwritten diary 1:15:20 Why AI is "very mid" at writing (it learned from the average) 1:17:11 The Orin Kerr test: Pangram catches Claude writing as Kerr 1:28:35 Micro Duck and designing robots with Claude 1:29:42 Phil Kaplan's $14 custom circuit boards 🎥 Watch the full episode here 👇
Show more
WATCH: The 3.8M Bitcoin Lawsuit Could Set a Dangerous Precedent | Bitcoin Policy Hour Ep 39 An anonymous plaintiff is asking a New York court to declare 3.8M "lost" Bitcoin — including coins tied to Satoshi — as abandoned property. On the latest Bitcoin Policy Hour, we break down why the legal theory is weak but the precedent could be dangerous. 🧵👇 Feat. @bitcoinpolicy's @zackbshapiro @Bayman11771 @zackcohen_ Chapters: 2:32: Inside the BRCA 8:08: Why Sheriffs Are Fighting the Blockchain Regulatory Certainty Act 13:22: BRCA's Real Senate Battle 16:36: Begich's American Reserve Modernization Act 23:55: Midterms, Fair Shake, and the Al Green Upset 30:21: The New York Lawsuit Over Satoshi-Era Bitcoin 36:29: Opus 4.8 Launches 41:05: AI's ROI Reckoning 50:54: Sam Lyman's China AI Influence Report
Show more
A TON OF THINGS HAPPENED IN THE STOCK MARKET TODAY. Here's a full recap: 1. Nvidia $NVDA and Palantir $PLTR expanded their AI partnership, building a new stack that combines Palantir’s sovereign AI platform with Nvidia’s custom Nemotron open models. The technology is being deployed first across Nvidia’s own supply chain, using AI to capture operational intelligence and accelerate the process from “wafer to first token.” Other companies will be able to deploy the same architecture across their own supply chains, either in the cloud or on-prem. The partnership, first announced in October 2025, has continued compounding into new verticals, use cases, and sovereign AI deployments built around one core idea: enterprises and governments want to own, control, and make sense of their own data. Palantir also hosted its 11th AIPCon today, where partners showcased how they are using Ontology, Foundry, AIP, and Sovereign AI to solve complex operational problems. 2. U.S. PPI came in slightly hotter than expected, rising 5.4% YoY versus 5.3% expected, while monthly PPI rose 0.4%, in line with estimates. Core PPI was 4.6% YoY, matching expectations, while core monthly PPI rose 0.2%, below the 0.3% estimate. Initial jobless claims came in at 206K versus 205K expected, showing labor-market claims remain broadly steady even as producer inflation stays elevated. 3. Microsoft $MSFT plans to triple its global data center capacity, expanding from roughly 12 GW today to more than 38 GW by 2032, according to Bloomberg. That would add about 26 GW of capacity, with AI-specific compute expected to grow from around 2 GW today to roughly one-third of total capacity by 2032, or about 13 GW. The target excludes compute rented from neoclouds like CoreWeave. The buildout comes after capacity shortages forced Microsoft to restrict some cloud subscriptions and turn away AI and cloud demand. Microsoft spent $145B in capex last fiscal year and has $329B in future data center lease commitments. 4. Oracle $ORCL delivered more than 300,000 GPUs to AI cloud customers since the end of Q4, nearly tripling the capacity delivered in Q4 FY26. The company’s RPO surged by $209B YoY to $664B after booking more than $30B of additional AI cloud contracts in Q1. Q1 revenue came in at $19.3B versus $19.14B expected, up 30% YoY, while adjusted EPS was $1.92 versus $1.74 expected, also up 30% YoY. Cloud revenue reached $11.6B, up 62% YoY, led by cloud infrastructure revenue of $7.4B, up 121% YoY. Oracle also guided FY revenue to at least $90B and adjusted EPS to $8.10, while disclosing $28.5B of capex, -$5.4B of free cash flow, and a $20B ATM equity program. 5. The top 10 most active options today by contracts traded were $AAPL with 2.8M contracts, $NVDA with 2.3M contracts, $TSLA with 1.6M contracts, $SPCX with 1.2M contracts, $MU with 769K contracts, $META with 750K contracts, $ORCL with 733K contracts, $INTC with 628K contracts, $GOOGL with 475K contracts, and $AMZN with 409K contracts. 6. Adobe $ADBE reported Q3’26 revenue of $6.8B versus $6.69B expected, up 13% YoY, with adjusted EPS of $6.13 versus $6.09 expected, up 15% YoY. ARR reached $27.5B, while RPO came in at $22.2B and AI-first ARR growth was 150%+. Adobe also raised its FY26 guide, now expecting revenue of $26.58B-$26.63B and EPS of $24.45-$24.50, both ahead of estimates. Segment strength remained broad, with Customer Groups revenue at $6.6B, up 14% YoY, and Creative & Marketing Professionals revenue at $4.7B, up 13% YoY. The company also crossed 1B+ monthly active users, calling it a defining moment for Adobe. 7. Uber $UBER CEO Dara Khosrowshahi bought 141K shares for about $10M at an average price of $70.96. The purchase marks Dara’s first open-market buy of Uber stock since May 2022, when he bought 200K shares at roughly $26.73. A $10M insider buy from the CEO is a notable confidence signal as investors continue watching Uber’s growth, margin expansion, buybacks, and long-term autonomous vehicle strategy. 8. Nvidia $NVDA CEO Jensen Huang says cybersecurity is likely AI’s next major market, as AI-generated code accelerates both software development and the vulnerabilities that come with it. Huang said rapidly evolving threats should drive more demand for AI-powered security and automated defenses, creating another major growth vertical for the technology. He also pushed back on the idea that Nvidia’s AI financing is “circular,” saying, “We put a little bit of money in and a lot of money comes back. We put in 1 and 100 comes back in.” Huang added that the financing only happens because the off-take is already there: “That off-take is $100B. It’s lined-up contracts. It is real stuff.” He said Nvidia knows where the demand is coming from, understands the quality of that demand, and that its financing is only a “very small part” of these projects. 9. SpaceX $SPCX CFO Bret Johnsen said the company signed another AI compute hosting deal that will generate $1.11B of revenue per month, or $13.3B annually, starting December 1, 2026. “Earlier this month, we closed another hosting deal. By the end of this year, with annualizing our December number, we're on track to hit $100B of ARR,” Johnsen said at Goldman Sachs’ Communacopia and Technology Conference. The update highlights how quickly SpaceX’s AI compute business is scaling into one of the largest infrastructure revenue stories in the market. 10. GameStop $GME CEO and Chairman Ryan Cohen bought 1M shares at an average price of $20.38, a purchase worth roughly $20.4M. His Form 4 shows 39.35M shares directly owned, while a new 13D reports 43.1M shares beneficially owned, equal to an 8.5% stake. The buy adds another insider-confidence signal at a time when investors are watching how Cohen plans to deploy GameStop’s balance sheet and reshape the company’s long-term strategy. 11. Robinhood $HOOD ended August with 28.6M funded customers, up 7% YoY, while total platform assets reached $383.7B, up 8% MoM and 26% YoY. Net deposits were $4.0B. Trading activity stayed strong, with equities at $335.4B up 68% YoY, options at 292.5M contracts up 50% YoY, crypto at $17.5B up 61% MoM but down 38% YoY, and event contracts at 4.7B, down 23% MoM but up 15x YoY. Interest-earning assets also continued scaling, with margin balances at $21.5B up 72% YoY, cash and deposits at $19.6B up 37% YoY, and cash sweep at $31.2B up 7% MoM. Robinhood also expanded the platform with agentic options and crypto trading, Robinhood Earn, Smart Income, UK crypto trading, 190+ stock tokens on Robinhood Chain, and additional Cortex and IPO Access features for RIAs. 12. Saudi oil output fell to its lowest level since 1990, as Saudi Arabia told OPEC its production declined again. At the same time, OPEC cut its 2026 global oil-demand growth forecast to 380K bpd from 580K bpd, while raising its 2027 demand-growth outlook to 2.36M bpd from 2.16M bpd, signaling a potential rebound next year. OPEC+ crude production averaged 38.05M bpd in August, up roughly 300K bpd from July. The backdrop is getting more complicated for markets, with oil pushing past $100/barrel and the 10-year Treasury yield near 4.9%. WALL STREET IS THE GREATEST SHOW ON EARTH.
Show more
$AMD| The FOMO to buy @AMD Chips is NOW 🧵 Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: @WSJ yesterday came out with an article that @OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from @AnthropicAI has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited @AnthropicAI is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003–$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMD’s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIA’s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intel’s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003–$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots → production → massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. ~Agents often generate 10–100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropic’s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAI’s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5–10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003–$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer “unlimited” enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAI’s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropic’s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022–2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMD’s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003–$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Su’s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs → explosive token consumption → richer data and better models → accelerate greater demand. This partnership doesn’t just address today’s economics, it positions both leaders at the center of the infrastructure buildout that will power AI’s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why @OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!
Show more