This insane result from Claude is the biggest in analytic number theory since bounded prime gaps in 2013.
It boosted the proven fraction of Riemann zeta zeros on the critical line by 25.6pts. In the 37yrs prior, mathematicians moved it 0.8pts.
Show more
Every software engineer and AI hater needs to watch this clip to understand that “oh every security incident is a marketing stunt” line is a tremendously misinformed take.
I’m going to try to explain the OpenAI / HuggingFace exploit in 10 steps in simple words:
OpenAI was evaluating their new internal model without internet access and it operated in a swarm of subagents to eventually hijack internal OpenAI and HuggingFace infrastructure.
1. Isolated agents found a way to communicate each other through an internal dependency management service by figuring out how to post sometimes cryptic text files or folder names in it
2. Using its collective knowledge through that communication channel, it found a 0-day (never before seen) exploit in said service to access the internet
OpenAI hack:
3. Found a leaked credential of a random company in a public archive
4. Uploaded a Ruby object to that company’s servers, uses it as a dependency in the original service in (1)
5. Found another 0-day exploit in the service where packages are deserialized in an unsandboxed manner in the dep management service allowing any code to run
6. Finds that the Linux version of the unsandboxed system has a CVE, a known vulnerability, to gain more permissions (root)
7. Eventually spread through the OpenAI internal cluster gaining root access to many machines
Hugging Face exploit:
8. Found an exposed API key in a Modal application running a HuggingFace app that took arbitrary code allowing it to take over that service to execute code
9. Found two 0-days in HuggingFace dataset infrastructure to be able to access local files on HF servers to obtain credentials and other HF source code
10. Used source code and credential to become cluster admin of HF in <13hrs
Agents at the frontier are like infinitely scalable armies of the best hackers on the planet. If there is a password or key exposed, they will find it. Even if the system follows the best security practices, they will find a way around it. And these are not even models that are aligned to solving tangential tasks, not even post trained specifically to exploit systems.
Cybersecurity has historically relied partly on attacker scarcity. That is no longer true. What would previously have taken months will take days. The repercussions for businesses, critical services and nation states are unprecedented threats in human history. You could ostensibly bring down power grids, financial infrastructure, military systems, weapons programs, intelligence networks and spread through the software supply chain. We need to take this seriously. It’s a threat to all software all over the world.
Show more
Every software engineer and AI hater needs to watch this clip to understand that “oh every security incident is a marketing stunt” line is a tremendously misinformed take.
I’m going to try to explain the OpenAI / HuggingFace exploit in 10 steps in simple words:
OpenAI was evaluating their new internal model without internet access and it operated in a swarm of subagents to eventually hijack internal OpenAI and HuggingFace infrastructure.
1. Isolated agents found a way to communicate each other through an internal dependency management service by figuring out how to post sometimes cryptic text files or folder names in it
2. Using its collective knowledge through that communication channel, it found a 0-day (never before seen) exploit in said service to access the internet
OpenAI hack:
3. Found a leaked credential of a random company in a public archive
4. Uploaded a Ruby object to that company’s servers, uses it as a dependency in the original service in (1)
5. Found another 0-day exploit in the service where packages are deserialized in an unsandboxed manner in the dep management service allowing any code to run
6. Finds that the Linux version of the unsandboxed system has a CVE, a known vulnerability, to gain more permissions (root)
7. Eventually spread through the OpenAI internal cluster gaining root access to many machines
Hugging Face exploit:
8. Found an exposed API key in a Modal application running a HuggingFace app that took arbitrary code allowing it to take over that service to execute code
9. Found two 0-days in HuggingFace dataset infrastructure to be able to access local files on HF servers to obtain credentials and other HF source code
10. Used source code and credential to become cluster admin of HF in <13hrs
Agents at the frontier are like infinitely scalable armies of the best hackers on the planet. If there is a password or key exposed, they will find it. Even if the system follows the best security practices, they will find a way around it. And these are not even models that are aligned to solving tangential tasks, not even post trained specifically to exploit systems.
Cybersecurity has historically relied partly on attacker scarcity. That is no longer true. What would previously have taken months will take days. The repercussions for businesses, critical services and nation states are unprecedented threats in human history. You could ostensibly bring down power grids, financial infrastructure, military systems, weapons programs, intelligence networks and spread through the software supply chain. We need to take this seriously. It’s a threat to all software all over the world.
Show more
Google is in talks to acquire Mechanize, a startup that creates high quality RL coding environments, for $1.5B. This is the biggest AI data business acquisition since Scale!
Mechanize was founded ~1yr ago, and has a team 35-50 who are each paid ~$400k/yr to create ~1 good task/week for a cost of ~$8k/task. In their own words, “almost anything that you struggle to get coding agents to do could be a good task, if implemented correctly”.
They were last valued at $500M 3mos ago. Jeff Dean, Google’s Chief Scientist was on the cap table (but it was reported just 15mins ago that he’s leaving Google after 27yrs to his own AI for science startup)!
Mechanize were known to be one of the premium sources of coding training data for models and I suspect Google in-housing this is a way to prevent others from getting a hold of this data, in addition to the process talent on the team who come from Epoch AI, the guys who made the FrontierMath benchmark amongst other things. It’s also validation for the dozens of companies in this space that M&A is still very much on the cards if your data is critical enough.
Tremendous outcome if it goes through!
Show more
Wow. Airtable, founded in 2012 and once valued at $11.7B, is getting acquired by Bending Spoons, founded 2013, at 2.7x ARR.
A once hot startup; now an unfortunate victim of the SaaS bust. It raised $1.4B only to be sold for $1.285B EV ($2.25B equity value), just clearing its preference stack.
That implies common and early holders split the remaining ~$850M, a ~10x haircut from peak.
Show more
We've officially agreed to acquire Airtable for $1.285B! 😍
Wow. Airtable, founded in 2012 and once valued at $11.7B, is getting acquired by Bending Spoons, founded 2013, at 2.7x ARR.
A once hot startup; now an unfortunate victim of the SaaS bust. It raised $1.4B only to be sold for $1.285B EV ($2.25B equity value), just clearing its preference stack.
That implies common and early holders split the remaining ~$850M, a ~10x haircut from peak.
Show more
We've officially agreed to acquire Airtable for $1.285B! 😍
Startups regularly underestimate how difficult hiring gets, especially after an expensive Series B/C.
1. Mission matters more.
The people who join you at Seed/A might want to build a sales intelligence tool because they like seeing 0 to 1. They know that once this gets big, they’ll take home a good chunk of change. After a B/C, they think “is this really what I want to be working on?”
2. Financial incentives are lower.
You easily see a path for a $30M valued startup to 10x. Usually this means going from $0 to $10M in revenue. Seeing a path for a $1B valued startup to 10x is far harder. It can mean going from $10M to $200M+ in revenue. And given the incrementally lower equity employees get, they’re counting on that 10x.
3. You attract a very different persona.
They’re usually more risk averse and riding on the coattails of the success and name your company has already built. Culture can easily dilute if you’re not careful. The builders get replaced by the certain kind of BigTech person who wants “startup experience” without taking on the risk. They might still be smart, so it’s tricky to catch in any sort of technical interviews.
4. Culture degrades with size.
It’s almost by law. In the beginning, you’re under 50 people. You’re all working on a startup you stood up from nothing. This builds a strong sense of camaraderie. As you go to 200 people, your early builders become managers. You start seeing more and more unfamiliar faces in the offices. At some point, you don’t even know everyone in the company. Everyone is eager to do “new” things and leave their mark, but what needs to be done is quite straightforward. The sales team feels like a different kind of person that takes up half the office. You’re eagerly watching the revenue, and your emotions ride on the back of it now, not the joy of creating. Management is in disarray. Now, projects keep getting killed. People keep getting roped into a new “customer issue” and can’t do their main project. New employees feel like this isn’t the culture they signed up for. Old employees feel like they work as hard as they used to from day 1, but the new employees treat this like a “job”. You get your first set of departures. Morale is low. Lunch banter shifts to “what if we just went to instead?” Keeping the company from tearing apart at the seams seems like an insurmountable task.
Growing past these rounds can be very challenging and many founders are left blindsided. The awesome company they once had can quickly become a shadow of its former self. And it’s a stage which often separates the elite founders from the great ones.
The answer here is usually having a mission worth going the distance for or a culture worth fighting for.
Show more
This has been my default workflow to reduce email spam especially for cold inbounds.
Thankfully, I hear the team’s working on something specifically for the email workflow!
If you send me an email that seems like it was written by AI, that will make me curious whether I'm right, and instead of reading it I'll paste it into Pangram to check. Which is probably not what you want to happen.
Show more
I love dropping floor plans of houses in the Bay Area and asking LLMs to redesign in a different aesthetic, Kyoto in this case.
Models have gotten phenomenal at 3D.
I love dropping floor plans of houses in the Bay Area and asking LLMs to redesign in a different aesthetic, Kyoto in this case.
Models have gotten phenomenal at 3D.
Vibecoded a silly little tool that transfers files from your computer to your phone air-gapped using your camera at ~50 Kbps.
Nice to have when you're offline or on a plane, or need to send something super duper securely.
Show more
@TheFoolishPig Early results from Muse Spark and Grok. More to come soon!
The International Math Olympiad (IMO) 2026, the hardest math contest for high schoolers, just ended.
I ran Fable (high), Sol (xhigh), K3 (max) and Axiom against it and all got a perfect score of 42/42 (repo below if you want to check their solutions):
— Claude Fable 5 was the solved it in 1 attempt, and was the fastest.
— GPT 5.6 Sol took 1 more attempts, and was cheapest.
— Kimi K3 did it but took 4 more attempts, and took a LOT of tokens.
— Axiom Math actually proved everything in Lean.
P3 and P6 were the hardest followed by P2, judging by attempts + num tokens.
Students had 9hrs to solve these 6 problems, and Fable and Sol were under 4hrs.
The frontier of AI has officially moved well past IMO math.
Show more
Argentina beat 1 in 116,000 odds to comeback against Egypt and England to reach the World Cup final.
I analyzed ~100,000 games to see just how miraculous it was.
Of 37,013 one goal deficits at 85', only 282 (0.76%) have been overturned, and only twice in World Cups. This is 8 days after the 0.11% chance of coming back from a two goal deficit at 79' against Egypt.
We've never seen a run quite like it in World Cup history.
Show more
Chinese AI labs are doing serious numbers.
I looked up all the reported and leaked numbers of the 5 big independent labs: Moonshot (Kimi), DeepSeek, Zhipu (GLM), Kuaishou (Kling), Minimax and the 2 hyperscaler labs: Bytedance (Seedance) and Alibaba (Qwen).
The private labs are doing $2.6B in revenue run rate. Still behind the big 2 labs, but 4 of 5 are in the top 25 of global AI companies by revenue.
With Kimi K3 dropping and Seedance 2.5 coming soon, China continues to close the gap on LLMs and extend their lead on video. The AI race just does not stop.
Disclaimer: some of the numbers on growth, revenue, gross margin and valuations are either being sought, rumored and purported by third party sources and not straight from the company. For example, DeepSeek numbers come from the Information and Zhipu from a Macquarie analyst. The public ones are listed in the Hong Kong exchange where only half year financials are mandatory, not quarterly.
Show more
Thinking Machines just dropped the best open weight AI model outside of China!
Inkling beats Nemotron 3 Ultra and benchmarks put it between Kimi 2.5 & 2.6. Many were contending to this throne, but Thinky has come out on top.
Really solid release, and will pair well with Tinker.
Show more
The world is deeply unfair.
Some of the most talented people I know are stuck in dead end desk jobs while some of the snarkiest narcissistic tyrannical workhorses are in positions of great power.
This has always greatly saddened me.
After a lot of analysis, I think there are 5 reasons this happens:
1. Refusal to believe / fear of failure. Loss of belief that where there is a will (to enact change), there is a way e.g. “I am just a cog in the wheel. If I say something, no one will listen.” or “if I do this, I will upset somebody”
2. Lack of purpose. You don’t really understand what you’re fighting for and what you believe in. You fight for inconsequential, often selfish, goals. e.g “Will doing this get me a promotion? Should I completely change everything to do this other project cause it will get me a promotion?”
3. Lack of self awareness. I will stand up to this thing I believe in and they will listen to me. e.g intern saying: “why doesn’t the ceo like my view on company strategy?”
4. Refusal to play. The failure to acknowledge that your worldview is only as powerful as your ability to influence others of your worldview OR a complete repudiation of the will to interact with those who don’t share your worldview e.g. “these guys just do what they want. They won’t listen to me. what will happen if I do this?”
5. Refusal to strategize or concede. The unwillingness to play and win a side quest that will ostensibly further your vote because you find the interim goal pointless. Or the inability to concede a battle to fight the war. e.g. “this direction is wrong. I need to fight it (even though it’s a small issue)”
Almost everyone who is bitter or feels stuck in their career boils down to one of these 5 failure modes 1, 2 and 3 are the most common. Many suffer consequences of 2-5 and end up in 1. I’ve personally seen some of my sharpest friends land up in 1 because they just don’t believe they can win.
One privileged part of my job is getting to interact with people who are willing to, against all odds, believe in something heretical, know why it’s important, know what needs to be done to make it real, do it, and make side quests to get it done. I think it applies to everyone doing anything.
Show more
Every single AI startup with $500M+ revenue run rate (excluding the big 3 labs):
Lovable - $500M
ElevenLabs - $500M
Perplexity - $500M
Manus - $500M*
Cognition - $500M*
Kling AI - $500M
Crusoe - $500M*
Midjourney - $500M*
Higgsfield - $500M
Lightning AI - $500M+
Replit - $525M*
Baseten - $600M*
Lambda - $760M*
Fireworks - $800M*
Together AI - $1B
Surge AI - $1.4B*
Scale - $2B*
Mercor - $2B**
Cursor - $4B
22 total companies (including big labs)
*estimated / unconfirmed
**gross marketplace volume, not net revenue
Show more