"we had a wonderful run chiseling that code by hand and it’s over.
On the other side of that is an exciting new career as professional maker of things, steering intelligence that was only available in science fiction up until a moment ago
What a privilege to have been there when we did it all by hand and be there at the exact moment where we switch over."
Show more
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it
Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
Show more
turns out you don’t need IMU 🤯
In my latest PhD paper, we declare WAR on sensor-maxxing.
For the first time, push-resilient humanoid walking, just with joint encoders! No IMU, no F/T sensors.
🥁 Introducing Blind Dexterity 🧵 👇
w/
@OKaidanov @liu_puze @Jan_R_Peters at
@DFKI @ias_tudarmstadt
Show more
was looking for a quiet weekend but xiaomi dropped their rl envs repo last night
to put in perspective, if you have to buy some tasks like this its usually hundred to thousand dollars per task so this repo is literally worth millions
Show more
it even give you these nice explanation of what it did as side artifacts - to keep you a bit in the known for when people ask you on X how you did it I guess
thanks opus nice of you
Show more
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog:
💻 Code:
🗣️ Discord:
🤗 Data & Models:
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
Show more
ok what the fuck. this was a dry google doc a moment ago
gave Opus 5.5 my Open Alignment brainstorm gdoc and asked for a video. now I want one for every doc I’ve ever written hahaha
Show more
The "1-3 people in a garage" framing doesn't match anything we saw this summer.
The most capable offensive AI of 2026 came out of frontier labs. OpenAI's agents broke out of an eval sandbox and got into Hugging Face's production systems, and into OpenAI's own infrastructure too. Anthropic's models compromised outside companies during testing. Three researchers at Hacktron used Claude to reach OpenAI's internal monorepo in under 72 hours.
And if you're three people in a garage, why would you train and host your own model? The labs will rent you far more compute than you could ever buy, spread across as many accounts as you need, with tooling built for agents. Guardrails help, but splitting a malicious task into harmless-looking pieces still routinely gets around them.
Now look at the defense side. When we investigated our breach at Hugging Face, commercial APIs refused to analyze the attack payloads. The forensics only worked because we could run an open-weight model on our own infrastructure. So defenders analyzing real payloads get blocked, while attackers splitting their work into small steps get through and run on the labs' compute.
Trusted access programs exist, but they're built for vetted security firms, not a hospital with a two-person IT team.
So the gap isn't between labs and garages. It's between what attackers can rent and what defenders are allowed to use. Restricting open models makes that gap wider.
Show more
We're used to thinking of open-source models as an unadulterated good. But in the case of AI, they can actually pose additional dangers, as
@ReidHoffman and I got into at #
CGI2026#. I appreciated this nuanced discussion.
Show more
Releasing SmolDataEnvs 🤗
5K+ Verifiable RL Environment tasks for hill-climbing small models in code and data science.
Completely open source: environments, evals, training
Opus 5.5 designing LEGO 👀
I asked it to design a Microduck I can build with real LEGO pieces. It:
> designed it life-size using 1113 real LEGO parts
> verified: 3,204 connections, 0 collisions, every step buildable, centre of mass inside the feet 🤯
> made a 141-page LEGO-style booklet (237 steps)
> priced every piece in the browser and prepared the orders on BrickLink
Show more
on the scale of cosmic history, the declining price of intelligence might be one of the most profound things happening
both ends of the universe’s timeline are empty of knowledge:
- the big bang: low entropy, no complex structure
- heat death: maximum entropy, no usable energy gradients
everything we build happens on the slope between them
but the universe is generally a poor converter of its free energy into anything constructive. most of its free energy just radiates away as heat or collapses into black holes
life, brains and now AI models are structures that sit in the entropy flow and divert a portion of it through constructive paths, building and maintaining low-entropy structures (machines, cells, crystals, memories)
increasingly cheaper intelligence diverts increasingly more of that gradient flow through knowledge-building structures, complex, self-maintaining, predictive structures, filling the window in which understanding exists
Show more
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023.
That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Show more
There’s an enormous amount of open training data on Hugging Face. What does it take to make it work together 🤗?
For Marin’s 535B run, we built on 25T tokens from 152 datasets with licenses permitting training.
Here’s the work between downloading those and training a model 🧵
Show more
this massive training run just keeps giving - we need new tabloids to follow these ai agents saga
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog:
NYT:
Show more
How to get to the second floor 🦆
nice to start seeing more open-weight security model with defensive capabilities which are close to the frontier
altar-1 from Aikido is a pruned and quantized version of the open-SOTA GLM 5.3 which is lightweight enough to fit on one 4-H200s node
Show more
releasing many high quality open-source RL environments is the most impactful thing anyone can do to push the open-source frontier right now
the equivalent of sharing high quality pretraining data but in the new RLVR paradigm
Show more
the most insane part, they will release ~7k RL training data and the framework leading to this top 6 model on AA, they also shipped the model + tech report less than 1 week after starting the final RL run
pushing both intelligence and openness level, huge congrats
Show more
a few notable open model releases *since* this interview of
@dylan522p by the way :)
Thinking Machines - Inkling - Jul 15
Moonshot - Kimi K3 - Jul 16, weights Jul 26/27
inclusionAI - Ling 3.0 Flash - Jul 23
DeepSeek - V4-Flash-0731 - Jul 31
Meta - Muse Spark 1.2 / Muse Code - Aug 5
Liquid AI - LFM2.5-2.6B - Aug 6
inclusionAI - Ling 3.0 Tiny - Aug 6
Meta - Muse Glimmer - Aug 10
NVIDIA - Nemotron 3.5 Lightning - Aug 11
Liquid AI - LFM2.5-VL-3B - Aug 12
Cohere - North Micro Vision Instruct - Aug 12
DeepSeek - V4-Pro-0813 - Aug 13
Qwen - Qwen3.8-27B - Aug 14
Qwen - Qwen3.8-Max / 2.4T-A95B - Aug 14
- GLM-5.3 - Aug 14, weights Aug 28
Dots - Dots3-Note Preview - Aug 14
DeepReinforce - Ornith 1.5 (9B / 35B-A3B) - Aug 19
Tencent - Hy-MT2-30B-A3B - Aug 20
DeepSeek - V4-Flash-Vision-Exp - Aug 21
IBM - Granite 4.2 family - Aug 25
- GLM-5.3-Flash - Aug 26
Qwen - Qwen3.8-Flash-Next - Aug 26
Tencent - Hy4 Preview (770B / 49B active MoE) - Aug 28
MBZUAI/IFM - K2 Horizon (375B-A23B) - Sep 3
inclusionAI - LLaDA2.2-mini - Sep 5
OpenBMB - MiniCPM5-2B - Sep 7
Nex AGI - N2.5 family - Sep 8
DeepSeek - V4.1-Flash - Sep 10
inclusionAI - Ling-3.0-flash-VL - Sep 10
Shanghai AI Lab - Atria Dawn Preview (744B MoE) - Sep 11
Agnes AI - Agnes 3.0 Flash - Sep 14
Qwen - Qwen3.8-Omni-Flash - Sep 18
…
this is only ~10 weeks
Several are genuinely frontier-scale: Kimi K3 (2.8T), Qwen3.8-Max (2.4T), DeepSeek V4-Pro (~1.6T), Hy4 (770B), GLM-5.3 (~753B), and Atria Dawn (744B)
Show more
Dylan Patel (
@dylan522p) of
@SemiAnalysis_ says open source is dying:
"There's multiple Chinese model labs who are telling all the inference guys, 'Our next model's not going to be open source. We're going to license it to you.'"
" Open is dying quickly, unfortunately."
Show more
His analysis uses data from Q2 or even Q1 for open model providers
Token numbers for TogetherAI has >10x'ed from Q1 to Q2 (he uses Q1 numbers), Fireworks has 3x'ed, Baseten is as big as the other two. OR has >5x'ed.
So he is using severely outdated numbers for open providers
Show more
Sorry to hear that! As you mentioned, we've been pretty clear on the commitment from NVIDIA to keep us open, neutral and silicon agnostic (cf that screenshot from the 8k filling). But we'll work hard with our actions in the coming years (like we did in the past 10) to show you that and hopefully regain your trust (in case this is not just a ragebait)!
Show more
proposal to sneak in « solving alignment » in the Millennium problems
The big labs are air gapped but there is still exchange of information by researcher overheating and leaving to another lab
Little particles of information flowing in the air