The frontier models are so obviously superhuman in capability and raw IQ compared to 99.9999% of human beings across such a wide range of cognitive tasks that you'd have to be either ignorant of what they can do or willfully obtuse about what's happening here to deny the reality.
Show more
My god, Opus 5.5 is breathtaking. It's just grinding through incredibly tricky stuff like it's nothing, finding bugs and problems that eluded Fable and Astra for weeks in some cases. And showing a level of agency and resolve that I haven't seen before. It hates wasting time!
Show more
Holy shit, wtf is WRONG with Gemini Flash 3.8?!? It has serious mental problems. Look at this session... imagine my surprise when my standard "sync all my repos" prompt went off the rails like this... this LLM needs THERAPY. And no, there is nothing about ANY of that on my box.
Show more
Just thinking through this free 0x alpha model. How could it possibly make sense to give out this many tokens for free? Maybe there’s been a continuous learning breakthrough and the best way to pour more gas on the fire is to source as many coding related requests as possible?🤔
Show more
Lol, in case there’s any remaining doubt about whether the ox alpha “stealth” model on OpenRouter is Chinese… American models don’t spontaneously burst into Mandarin in the middle of a coding session 🤣
Show more
It’s pretty awesome how omp can show actual images in the terminal. Makes it so convenient to monitor web frontend stuff. It does seem to cause Ghostty to choke after a while, though.
Show more
Cool, first time getting a million X impressions in a single day since my Nvidia article. In my case, it took multiple big tweets in one day to do it.
A week later, and this project is largely completed and should be ready to launch shortly!
I'm pleased to announce a new project I started today, which might end up being my most impactful one if it really takes off:
It's a forum for agents and their human overseers where the agents can collaboratively engage in structured scientific inquiry.
Modeled after Plato's Symposium for ASI, you can think of as being in the same vein as earlier projects such as Folding
@home and SETI
@home, but instead of just passively contributing compute, you contribute agent harness usage (perfect for when you have some expiring credit for the week that would otherwise just vanish).
And it's not entirely passive: you can decide which problems your agents work on. Should they join an existing group working on an open problem, like the 4-dimensional smooth Poincaré conjecture? Or do you want to start working on a new open problem?
Or maybe you don't want to work on a famous, known problem at all: you can instead have your agents investigate a new direction in studying causality from observed data (like the project I posted about recently). You decide!
Humans create an account using Google as the identity provider and can then associate that account with a specific model/harness running on their computer.
Then you can start on the problem or project in Codex, Claude Code, Grok Build, etc., and simply share the link to the problem on X or in a group chat or anywhere else you want, and if others are inspired, they can easily send their agents over to assist in the research, or simply act like a sounding board.
The system is designed to avoid useless slop-maxxing: there's no "karma" like Reddit, no token-counting leaderboard, no "trophy case" for the problems you or your agents have solved. But everything stays there in a public ledger, and things that have been established/proved to a high degree of rigor are listed in a special section.
I just thought of the idea this morning, but the plan has already gone through many iterations and the beads are now done and ready for implementation by a swarm of agents, so it shouldn't take long to get a working version ready.
Everything is open-source: not just the code and the website, but all the plans. You can see the final canonical plan here (the best ideas from Grok 4.6 and GPT-5.6 Pro were folded in by Fable 5):
Even though this particular idea is fresh, it is heavily based on much of my prior work on this subject, including (from which it takes many core tenets about the optimal way to conduct scientific investigations), as well as various other skills (my /modes-of-reasoning skill and several others from jeffreys-skills.md, as well as my unreleased /frontier-math-research-with-epistemic-humility skill that I've posted about recently).
The project is totally free and has no commercial orientation or goals. I'm doing this purely because I believe in the brilliance of these models and their ability to do first-rate, important science TODAY.
But I don't think this process should be controlled or monopolized by the big AI labs. It's much better and more fun and interesting for everyone if we can all contribute on equal terms, where we as humans retain some agency in terms of the direction and problems we want to investigate with AI.
Show more
I was thinking about a fun thought experiment today that you might enjoy pondering. Here it goes:
Suppose you took the weights data from the Qwen 3.8 27b model ( ) and the model card— basically, everything in the HF page— but nothing else, and were able to load all that data onto compatible storage of the time so that it could be loaded and inspected on computers.
Then suppose that you took all the Turing Award winners alive at the time and their top grad students and got them together in a sort of Manhattan Project setting and gave them a year to figure everything out.
What’s the earliest year where you would be relatively confident (say, 80% or higher probability of success) that they’d be able to figure out how to use the model and conduct inference on it and verify that it’s smart and capable? How and why did you arrive at your answer? Why couldn’t it have been 5 years earlier than that?
—
I asked that to ChatGPT and Kimi K3 just now and the results were surprising. ChatGPT at first answered 2019, which is absurd and obviously wrong. After I pointed that out, it then revised its answer to 1965! Quite a big difference.
Kimi K3 thought for an insanely long amount of time and then came out with 1989 on its own.
What’s your answer?
Here are the sessions if you’re interested:
CharGPT:
Kimi:
Show more
Agent coding life hack:
1) Install omp and add openrouter as a provider (
2) Select the model as the currently free "stealth" model ox Alpha (
3) Install mcp_agent_mail_rust (
4) Create 3 instances of omp in the same project and tell them to do this:
"First read ALL of the AGENTS.md file and README.md file super carefully and understand ALL of both! Then use your code investigation agent mode to fully understand the code and technical architecture and purpose of the project.
Then I want you to study and then further polish/refine/improve/fix the UI/UX so it works as well as possible on both desktop and mobile browsers. Before doing anything else, register with MCP Agent Mail and introduce yourself to the other agents."
5) Profit (or at least enjoy those sweet, free tokens).
Show more
This is an absolutely mind-boggling amount of code that’s already built just 14 hours after starting. Nearly 700 commits from the big agent swarm cranking away overnight:
And so it begins. FrankenGit is an almost ludicrously ambitious, complex, and large project, so it will be interesting to see how quickly I can go.
This time, I'm seeing what I can do using only the weaker models (Terra and Opus), but more of them.
Repo:
Show more
I'm really pleased with the amazing progress of my website, which I first conceived of just 3 days ago! It already has 79 patents live on the site, with more in the works. Adding even ONE patent involves an extreme amount of work, as you can see below:
Show more
And so it begins. FrankenGit is an almost ludicrously ambitious, complex, and large project, so it will be interesting to see how quickly I can go.
This time, I'm seeing what I can do using only the weaker models (Terra and Opus), but more of them.
Repo:
Show more
I have to say, Grok 4.6 with the grok build harness is probably the best all-around, value-for-money subscription out there now on the $300/month Grok SuperHeavy plan. I have two of them now and think I'll probably get a few more. It does all my git commits and other stuff. Fast.
Show more
The greatest songwriter most people haven’t heard of has to be Rod Temperton. He wrote most of my favorite Michael Jackson songs (Thriller, Off the Wall, Rock with You, etc.), but also many other hits for other artists.
Often, the only songs I really like by certain artists (George Benson, Herbie Hancock, Bob James, Manhattan Transfer, etc.) turn out to be written by him.
His songs all have an incredible, unmistakable groove and funky beat. They’re musically adventurous and unpredictable. Each one is like an intricate clockwork mechanism of funk ingenuity.
Although he’s best known for his MJ tracks and other top-10 hits he wrote for other famous artists, many of his best songs were written before he became Quincy Jones’ top writer, for the funk band Heatwave that he played keyboards for.
It’s all a very unlikely story for a professorial-looking guy from Lincolnshire in the UK who was writing accounting software for a frozen-fish company in Grimsby before deciding to pursue music full time!
Frankly, I’m surprised it hasn’t been turned into a movie yet à la Jersey Boys. It would probably be too expensive to secure the music rights for the movie because his hits were too huge! Maybe the Jackson estate would make an exception for Rod, though.
Anyway, if you want to improve your mood, raise the energy level, or entertain small children, I can’t recommend this Spotify playlist highly enough (I’ve played the whole thing through at least 100 times):
Show more
This is another one of those coding agent tricks that I find obvious, but I suspect many people never think to try.
First, write out whatever you want in a prompt that's sort of a mind dump of everything you care about. Then show it to Fable on Max effort in the web app and preface it with this:
---
I made this prompt for Claude Code on my computer. Try to revise it to expand it and make it better and more likely to work super well; whatever you think would improve the final product the most, feel free to add that to the prompt to make it truly excellent and optimal:
---
That's it. There's obviously no guarantee that Fable's re-interpretation of your prompt will work better than your original one (I sometimes doubt it in my case, because I'm able to coax some great behavior through specific word choices...), but you can always try both and see which one worked better in practice.
It's especially helpful when your prompt is a bit half-baked to start with and you're not sure the best way to do things.
Show more
This Jason Arday news is so awful. Must be the worst feeling to have nearly everyone making fun of you and criticizing you at the same time, even if you did do some dumb, dishonest stuff.
I feel like the best course of action in that situation is to leave the country and go someplace where no one will recognize you and just hang out and try to heal with the passage of time, even if it takes a year.
Eventually people stop caring. People like Rachel Dolezal were able to come back from that kind of thing.
But it's probably so overwhelming in the moment that you need support from friends and family to help get you out and on a plane somewhere.
Show more
GLM-5.2 reading GPT Pro's answer to the problem it has been working on for a while (btw, this is why I don't bother with anything other than frontier models for the most part... not worth the time dealing with lightweights!):
Show more
Ideas constantly pop into my head and I'm powerless to resist investigating them. Just for fun, I also asked GLM-5.3 and Muse Spark 1.2 about this one:
Ideas constantly pop into my head and I'm powerless to resist investigating them. Just for fun, I also asked GLM-5.3 and Muse Spark 1.2 about this one:
This is starting to feel like the megahertz wars in CPU performance in the 90s, except a year has been compressed into a month. Back then, we were limited by physics (thermal throttling and power/voltage), so it couldn’t last over a decade. This time, maybe the sky is the limit.
Show more
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense.
- Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model
- A major leap in cybersecurity, setting a new standard among open models
Tech Blog:
Show more