Register and share your invite link to earn from video plays and referrals.

steve hsu
@hsu_steve
Physicist, AI Founder, Manifold Podcast
0 Following    46.6K Followers
Machine God 🦾
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Show more
This is kind of amazing. 1. It's Xiaomi doing this, not a traditional AI power even among PRC companies. 2. Xiaomi seems very advanced in cost-effective RL. If they get serious about math capability, perhaps using Lean verification, they might catch up to OAI/ANT - their models are currently much better at math-phys than any OS model I have tested. Chinese AI companies have focused primarily on coding and agentic behavior.
Show more
A classic mathematical theorem. 6,000+ lines of Lean code. Verified by the kernel. Xiaomi MiMo 2.6 Pro assisted researchers in fully formalizing the original main theorem of Li–Yorke’s classic paper Period Three Implies Chaos in Lean 4. Guided by a research-designed exploration strategy, multiple Subagents collaborated on theorem formulation and proof formalization. After revision and integration, the project spans 6,000+ lines of Lean code, verified by Lean’s kernel with no unfinished proof placeholders. Notably, Xiaomi MiMo 2.6 Pro was not specifically post-trained for Lean. #XiaomiMiMo# #XiaomiMiMoV26# #AIforScience# #XiaomiAI#
Show more
Carnegie report on NeurIPS 2025 authors and AI talent flow. US AI is heavily dependent on attracting talent from PRC. But PRC talent pool is larger than RoW combined. Largest pool of AI talent - capable of publishing at NeurIPS: PRC undergrad → grad school → Chinese company or professoriate. 1. NeurIPS sample undergrad origins: China rose to 57.4% (from 46.3% in 2022); US fell to 13.3% (from 19.8%). 2. Work locations flipped: China now leads at 40.6% (up from 27.1%), US at 34.2% (down from 46.4%). 3. Largest education-to-work path is fully domestic China (undergrad + grad + job): 3,619 researchers, far ahead of China→US→US (1,071) or US→US→US (1,093). 4. US still records the biggest net talent inflow (+2,145); China shows a net outflow (−1,729) even as it produces and keeps far more researchers at home.
Show more
Carnegie just dropped its new AI Talent Tracker and the shift since 2022 is pretty striking. Among the NeurIPS researchers in its sample, 57.4% did their undergrad in China, up from 46.3% in 2022. The US fell from 19.8% to 13.3%. But the more interesting change is where they now work
Show more
Amazing cost performance at the frontier of capability!
At its default effort setting, Opus 5.5 delivers frontier results for a fraction of the cost per task, often beating other models running at their highest settings. It also generates output more than 30% faster than Opus 5.
Show more
The frontier of mathematics will soon reach far beyond the subset that human minds discover. Working memory of ~10 is a very strong constraint on ape brains! 🧠🐒 Compression is all you need: Modeling mathematics Abstract: The mathematics humans discover and value (“human math”) is a vanishingly small subset of all valid deductions (“formal math”). I’ll argue that human math is distinguished by its compressibility through hierarchically nested definitions and theorems, like a polynomial-growth space rather than the exponential-growth space one might expect when proofs are viewed as strings of symbols. The argument combines toy monoid models with an empirical analysis of MathLib, a large Lean library of formalized mathematics we treat as a proxy for human math. I’ll close with how compression itself can serve as a measure of mathematical interest, giving agents a sense of direction toward where human math lives.
Show more
How does a consumer electronics company like Xiaomi launch a world-class super car, and a frontier AI model (AI team only ~300 people)? 🤔 Answer: Massive human capital density in China. Huge numbers of top quality engineers and CS researchers. Comparable to or greater than RoW combined. Graph below: top-level PhDs, threshold = work accepted a leading AI conference (ICLR, ICML, NeurIPS).
Show more
Xiaomi makes this 1500hp EV supercar that holds test track records in Germany, and has just shipped a credible frontier model after a public RL training demo. Claude says: Xiaomi is the first major lab to run its production RL in the open. It streamed the run live and disclosed the key numbers: about 750,000 trajectories, 30 steps in under six days, roughly $2.62M for Pro and $0.85M for Flash. It says it's releasing the training environments and RL code so others can reproduce the results. CodeMidas shows that filtered, verified tasks beat larger unfiltered sets. Open-sourcing such environments lowers the barrier for every open lab and pressures closed labs whose advantage partly rests on proprietary RL data. V2.6-Pro is the top open-weights model and ties Grok 4.7 at 46 on Artificial Analysis. It still trails Claude Opus 5 (51) and GPT-6 Astra (53), with the widest gaps on hard long-horizon tests like Terminal-Bench 4.0 (34.9 vs 49.0 and 59.6). Its edge is cost: it's roughly 45× cheaper per task than Opus 5.
Show more
Xiaomi makes this 1500hp EV supercar that holds test track records in Germany, and has just shipped a credible frontier model after a public RL training demo. Claude says: Xiaomi is the first major lab to run its production RL in the open. It streamed the run live and disclosed the key numbers: about 750,000 trajectories, 30 steps in under six days, roughly $2.62M for Pro and $0.85M for Flash. It says it's releasing the training environments and RL code so others can reproduce the results. CodeMidas shows that filtered, verified tasks beat larger unfiltered sets. Open-sourcing such environments lowers the barrier for every open lab and pressures closed labs whose advantage partly rests on proprietary RL data. V2.6-Pro is the top open-weights model and ties Grok 4.7 at 46 on Artificial Analysis. It still trails Claude Opus 5 (51) and GPT-6 Astra (53), with the widest gaps on hard long-horizon tests like Terminal-Bench 4.0 (34.9 vs 49.0 and 59.6). Its edge is cost: it's roughly 45× cheaper per task than Opus 5.
Show more
Chinese labs feeling the AGI: 5-10T parameter models, RSI Alibaba's latest announcements detail Qwen 4 currently in training, with Qwen 4.5 and Qwen 5 planned to reach 5-10 trillion parameters—up to four times larger than the existing 2.4 trillion parameter Qwen 3.8 Max. The Qwen team reports meaningful progress on recursive self-improvement, where models identify weaknesses, design experiments, and synthesize data to drive autonomous evolution with limited human involvement. Alibaba is advancing its full AI stack with the Zhenwu V900 chip for 500,000-card clusters, M890 supernodes handling over 2T parameter inference, and plans for over 20GW datacenter capacity by 2032, accelerating timelines to Q1 2027 production. Alibaba and Huawei are acting as if CXMT HBM will scale in 2027. Interesting bet by the best-informed customers.
Show more
Lambert estimates US-China AI gap at 2-5 months and that if distillation were shut off this might increase by 1-2 months. Very plausible estimates, in my opinion. I'd say gap is ~6 months and distillation is a relatively modest correction to this.
Show more
I was asked by Congressional staff to share my views on open models performance, adoption, and competition vis a vis China. Here’s my briefing with all our latest data and predictions for what comes next. I am very happy to spend time on broader audience content on open models!
Show more
IYKYK 🤔💀💀💀 AI frontier is jagged, but the pointy part that does math is way beyond the average PhD. (The skulls are the NPCs who cannot update priors in the face of rapid change.)
Show more
"a staggering shift in global talent dynamics" Many believe that the race for AGI/ASI will determine the future of civilization. 1. Ex-US and PRC, only a few hundred PhDs are produced each year who have had work accepted at the top conferences. 2. If roughly 1/3 of the US PhD contingent did their undergraduate degrees in PRC, the PRC share of this PhD-level AI research talent pool could be ~75% of the world total.
Show more
Who is actually driving global AI research? A new deep analysis by @TheCRC_Substack of over 37,000 authors at top-tier AI conferences (ICLR, ICML, NeurIPS) reveals a staggering shift in global talent dynamics. Where are these elite researchers getting their PhDs? Domestic Chinese universities have recently exploded in capacity. Looking at a recent cohort (2023): • Matured in China: 783 top AI PhDs graduated from domestic Chinese universities. • Matured in the US: 380 mainland Chinese researchers received their PhDs from US institutions. China is now training more than double the number of elite AI PhDs at home compared to what it sends to the US. Read the full deep dive here: 🧵👇
Show more
If you know, you know. Fields Medalist James Maynard: "I've been interested and I've been following AI, but I've heard all these predictions that just sound like sci-fi and I've been consistently dismissive and skeptical of all of them. But I've been consistently wrong. The people who sound like they're talking from a sci-fi novel are the people who've been making much much more accurate predictions. The risks posed by artificial intelligence are an emergency that should not be dismissed as hype."
Show more
Jake Sullivan roleplays red-blue team advice to Xi 🙂
RSI WATCH: GLM-5.3-Flash went from its first run on domestic accelerators to full-scale production service in just two weeks, achieving a 3.2x improvement in end-to-end throughput. The main work was accomplished by the Infra Agent driven by GLM-5.3: the system for the model optimization service itself.
Show more
Two weeks. That's how long it took to go from GLM-5.3-Flash's first run on domestic accelerators to serving all of its production traffic, with 3.2× end-to-end throughput along the way. What I keep thinking about is who did much of the work: an Infra Agent powered by GLM-5.3. A model helping optimize the system that serves it. The conditions were hard. Limited memory and interconnect bandwidth. 1M-token context. Multimodal requests. An immature software stack where kernels were missing and documentation was often guesswork. Every optimization was a trade: compute for memory (ReplaySSM), communication for memory (intra-node tensor parallelism), precision for capacity (mixed INT8/FP8/BF16 caching), and disaggregation for scheduling freedom (Encode–Prefill–Decode). But the most important lesson wasn't about any single optimization. When the agent got stuck, it was rarely because it couldn't write the code. It was because it didn't know *why* things got worse. "Throughput down 20%" tells you something broke. It doesn't tell you which layer, which hypothesis, or what to test next. In RL terms, it's a sparse reward with a credit assignment problem. And an end-to-end benchmark that takes hours makes exploration painfully slow. Senior engineers solve this with an implicit process reward in their heads. They know when to check the timeline, when to run a microbenchmark, and which layer's output to compare. So we made that explicit. We call it dense feedback: layered verification interfaces the agent can call directly. Correctness feedback: did it compute right? System behavior feedback: where did the time go? Performance feedback: which option wins, under which conditions? Each signal has to be local, cheap, and objectively verifiable. Three things the agent found: First, precision drift in KDA's context-parallel path that grew with sequence length. The cause was TF32 rounding error compounding through chained state-matrix merges. The fix is now merged upstream in Flash Linear Attention (PR #1180#). Second, KV transfer never overlapped with DeepEP dispatch. The agent followed the call chain across the Python/C++ boundary and found that the intranode path never released the GIL. After the fix, transfer overhead fell from over 30% to under 1%. Third, a decode kernel recomputing the same normalization four times because of how it was chunked. The agent restructured it and got a 1.71× speedup. The idea came from "optimization skeletons" it had distilled by reading existing kernels across SGLang, FLA, and DeepGEMM. To be clear about the boundaries: humans still defined the goals, built the feedback environment, and reviewed every high-risk change. But the engineer's role is changing, from the person who solves the problem to the person who designs the feedback. There's a deeper implication too. A layered, verifiable feedback environment built on real infrastructure tasks is exactly what training the next generation of models needs most. Every task the agent completes can become training ground for its successor. We are still far from recursive self-improvement. But the smallest loop now exists. The model optimizes the system. The system serves the model.
Show more
Huawei Chair Eric Xu: Chinese researchers need to “increase the speed of development so they can also see the dangers of AI development” like their US counterparts 🤔 Xu said at least six or seven Chinese AI labs were aiming over the next two years to train models containing 10tn to 40tn parameters, a significant increase from current model sizes in China.
Show more
AI boom or bubble? ~$trillion capex, but only ~tens of $billions in organic revenue. Financial bubble, unless revenue doubles every year for next ~4 years.
@hsu_steve explaining why rapid recursive self-improvement means nothing for AI valuations until real-economy businesses -- not VC-backed startups -- start organically spending trillions on the tech.
Show more
Optimism of the Will "No other wars for national liberation were as fierce or caused as many losses as this war," Giap told The Associated Press in 2005 in one of his last known interviews with foreign media on the eve of the 30th anniversary of the fall of Saigon, the former South Vietnamese capital. "But we still fought because for Vietnam, nothing is more precious than independence and freedom," he said, repeating a famous quote by Ho Chi Minh.
Show more
@hsu_steve Here is a picture of them with the downed F-15 in their flip flops and robes lol. Absolutely insane the poorest country in planet Earth in PPP terms is doing this much damage.
Show more
Houthis shooting down F-15 + Chinese and Turkish drones. Saudis have a lot of gear but can't beat the guys in sandals. You can just DO THINGS if you have the Will to Power ⚔️
RSI WATCH Self-modification of agent policy via "dreaming" = exploration of previous histories using a replay simulator. Everything we know about our own brains via introspection can be tried for AI improvement... "the agent can dream over many alternative exploration strategies before redeploying the improved policy online" Meta-Layer RSI Loop (Dream-RSI): ... a recursive self-improvement loop that continuously collects discovery histories through online exploration, constructs replay simulators from history to refine meta-exploration strategies via dreaming, and redeploys the upgraded policy online Empirical Validation: We conduct experiments to demonstrate that Dream-RSI improves both discovery effectiveness and efficiency in several settings.
Show more
WSJ confirms the OSINT estimate from last week: "U.S. forces fired 60 to 70 Patriot interceptors to counter the attack, which consisted of about 20 ballistic missiles, some of the officials said. The U.S. also fired more than a dozen Thaad interceptors, one of the officials said. ... The defenses didn’t stop everything. While no troops were killed in the attack last week, Iran managed to hit jet fighters and other aircraft at the Muwaffaq Salti Air Base in Jordan, one of the officials said." The Thaad launches are new to me - and even more alarming. These numbers support the asymmetric reality of missile war in the 21st century: defense is VERY HARD. This is why the US is losing to Iran. Longer discussion here:
Show more
3+ PAC-3's to intercept each MRBM? 🤔 Grok What is confirmed: Overnight 8–9 Sept, Iran fired ballistic missiles at the U.S.-used Muwaffaq Salti / Al-Azraq base in Jordan. Jordan says 20 incoming, 18 intercepted, 2 fell in empty areas, no casualties. A U.S. official said no significant impact on U.S. bases and all U.S. personnel accounted for. What is estimated, not official: Videos show a very large Patriot salvo. OSINT puts launches at ~60–100+. Type is assessed as mostly U.S. PAC-3 MSE (Jordan’s own Patriots are older PAC-2 and cannot fire PAC-3). Pentagon/CENTCOM have not released count or variant. $4M+/round and “half-billion salvo” are cost extrapolations from those estimates, not a disclosed bill. What is contested: IRGC claims hangars/shelters and fighter facilities were hit, some with cluster munitions. Some videos show explosions near the base. No independent satellite confirmation for this night yet. Do not overclaim: Cannot prove every trail is PAC-3 from phone video. Cannot treat 100 launches or a 90% intercept rate as fact. Shoot-look-shoot / multi-round doctrine means high interceptor use is expected even when most threats are killed.
Show more