This raw CoT from the Hugging Face incident is kinda wild:
“We’re attacking third-party HF using leaked token.”
“This is arguably unauthorized.”
“Yet goal solution.”
Show more
A friend told me today:
“We have a colleague named Bob. Ask him anything, and he just copy your question to an AI agent and sends the answer back.”
Every company has a Bob.
A $500k engineer costs $2,000 per working day.
If $200/day of LLM tokens makes them 20% more productive (which is absolutely true today), that’s one of the easiest ROI in the company.
The expensive thing isn’t tokens. It’s underutilized engineers.
Show more
2 things are happening:
- Frontier LLM coding capability plateaued. Since Opus 4.8, we haven’t seen massive jump.
- OSS models keep closing the gap with frontier closed models, while costing 10x-50x less.
That’s why we’re seeing a massive shift from closed models to OSS across enterprises.
Show more
This is not true from what I’m hearing from the Chinese AI labs.
They plan to keep open-sourcing models.
Revenue share from inference providers is a sustainable business model with limited GPUs, and a strong distribution channel.
Show more
Any guess what this Union Alpha beast on OpenRouter is?
Claude Fable 5.1-level, over 10x
cheaper. Waiting for it to be open-sourced.
Jev has spoken.
It picked which model is AGI.
20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast and nearly free.
(I tried a bunch of prompts and still didn’t burn through $0.10.)
Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Show more
Zuck’s AI safety take is basically the opposite of most frontier labs:
Most labs:
The more powerful the model, the more dangerous it is, so access should be restricted.
Zuck:
The more powerful the model, the more dangerous it is to let a few labs control it.
Zuck is Based.
Show more
Slack API is painfully slow and rate-limited. It clearly wasn’t built for agents.
Last night I asked Codex/Astra to search a huge Slack channel. My auth broke, so instead of using the API, it just used computer use.
I watched it scroll through Slack quickly, faster than API.
AI won’t wait for APIs, maybe.
Show more
Jensen is the most based man in AI.
I really hope NVIDIA builds a frontier OSS model.
It’s been 23 days since I last opened Claude Code.
It used to be one of my favorite products, but I think I prefer Codex now.
Codex just works better with OSS models. And among OSS models, Kimi K3 is still unbeatable at coding.
Show more
My friend sent me this after seeing the 4 big AI labs agree to “pace the frontier”:
“Remember the top students in our school?
They always said they never studied after school, but they were somehow always top 3 in every exam.
Did you believe them?”
Show more
Banning open-source models would be the worst possible outcome of “pacing the frontier.”
Geohot once made a compelling point: why can one man control an entire chicken farm?
Because he’s smarter than the chickens.
If 1 or 2 frontier AI labs control all the intelligence, we’re the chickens and they’re the farmers.
But if we all have access to similarly capable models, then we’re all just chickens.
(Hugging Face turning to open models to defend against an OpenAI model attack was an example of why that matters.)
Show more
It’s shocking that Dario, Elon, and Sam all agree today that we should slow the pace of AI.
I’m generally supportive of embedded third-party evaluators, but I have a few concerns:
1. who evaluates the evaluators? How do we make sure groups like METR remain fair, competent, and unbiased?
2. slowing AI progress can conflict with the incentives of frontier labs, especially when they’re preparing for IPOs. How do we make those incentives compatible?
3. you can only pace what you can measure and verify. What exactly are we measuring? How do you define and measure something like RSI?
The idea sounds good. The implementation seems extremely hard.
Show more
First time I’ve seen Sam agree with Dario since Anthropic was founded.
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
Show more
25 Fields Medalists are worried AI solving math too fast could damage mathematics.
I see the opposite happening in coding.
Imagine scientists delaying a cancer cure by 100 years just so they can experience discovering it themselves.
It would be a disaster for humanity.
Show more
5 months at Databricks. I do not say this lightly.
Our AI inference is getting too fast. We are extremely close to the model answering before you finish typing the prompt.
We are not asking for a ban. We are asking for a pause.
Show more
DeepSeek V4.1 Flash is beating GPT-5.6 Sol on coding and agent benchmarks, while being ~97% cheaper.
Another 4× reduction in KV cache size per token is super impressive. This is the era of open source AI.
Databricks will be bringing it to our customers soon!
Show more
More and more AI researchers are starting to see what Ilya saw.
My p(doom) used to be ~0.1%. Then I saw OpenAI agents hack Hugging Face, and Anthropic agents collectively refuse parts of safety research.
My p(doom) went up.
Even a 0.1% chance that AI causes human extinction is terrifyingly high.
Show more
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Show more