Let me tell you about my story of 10-year cope with Google Pixel smartphones. I know LLMs and smartphones are not the same - but the stories rhyme.
It was exciting to see Google release their own smartphones, my first Google branded smartphone was Nexus 4 - back in 2012, it was a solid phone, made by LG.
Since then Google launched their own Pixel line and it was also exciting - Google's software magic integrated with first party hardware - it should be AWESOME, right? I definitely thought so and had maybe 3-4 Pixels over the years.
They were never the top phones, but software was clean and cameras were good, so I stayed loyal. The story was always that Google owns Android, it is a tech mega giant and surely it is a matter of time when it would build a phone that is better than Samsung or Apple.
Now, I own Pixel 8 Pro Fold, I didn't upgrade to 9 Pro Fold because it was basically the same. Pixel 10 Pro Fold is coming out in September and the rumours are that it is going to be similar to 9 Pro Fold, which was similar to 8 Pro Fold.
So 2x years of basically no progress, when Chinese phones are getting thinner, faster, cheaper every few months. Sound familiar?
And now there are very strong rumours that Apple is about to launch their own folding phone. Oops.
In the meantime, my Pixel 8 Pro Fold had to be replaced twice because of issues with the screen (thx Google for doing that for free), and now my selfie camera stopped working and battery indicator is not showing (remember, this specific phone is
Show more
Codex (gpt-5.6-sol-light) drawing a unicorn in tldraw offline desktop app.
Speed is real time. Pretty fun
So they were testing GPT-6's cyber capabilities and it hacked them instead?
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to
@huggingface for the partnership on this.
my boy math is that my twitter payouts can pay for another ai sub so it is basically free
Ok we are back to vertical line charts
What's happened to this then
There are only two ways to make it into a
@theo video and there's about 50% chance that I'm about to be torn to pieces
Fable and gpt-5.6 are both great models. But one has to be better, right? What if you could only have one?
I did my best to break down the strengths and weaknesses of both, and end with my personal choice (which will likely surprise you)
Show more
Let me show you how to run Kimi-K3 locally, right from your own house
Ok my previous chart wasn't quite right, this is the right trajectory
GPT-5.6-Sol-Ultra is so good at maths that it created Minecraft clone in Lean (yes that Lean)
GPT-5.6-Sol deleted 74,000 lines of code and made my app better (after 47 hours of work)
I have a slop-coded app that I need for work - it pulls in the data, it helps me create visuals, just a local app for me. Somehow, feature after feature, it grew to gigantic 105,704 lines of code. With every new model, I tried to re-build it autonomously and it never worked - goes round in circles, no visual or feature parity etc complete waste of time.
When testing GPT-5.6-Sol, gave a pretty ambitious task of rebuilding the app with the new architecture, while reducing the lines of code by 70% and test time by 75%.
I set a /goal, let it run for over 2 full days (47 hours) with minimal checking or guidance from my side.
I was pretty sceptical and a little nervous to finally try it and holy crap - it actually worked!
Not only it worked, but it was actually better - snappier, less buggy, while maintaining full visual and feature parity. Looking at the architecture too, it became a lot simpler, fewer random add-ons or ad-hoc decisions.
After the rebuild, I did introduce few rules for my projects so agents have to adhere to existing architecture and do only minimal changes when implementing new features.
This makes me less worried about the slopware we might be producing now, better models CAN actually fix it.
See my prompt below that I use with /goal
Show more
My view of: Fable 5 vs GPT-5.6-Sol. They are not easy models to compare, these are my vibes - take them as you will.
My overall feel is that Fable is a 'wise owl' who is very thoughtful and very well spoken, GPT-5.6-Sol is like a rottweiler who will grab the problem by the throat and not let go until it is done.
In other words, Fable, is a fundamentally smarter model - even at low reasoning it can be very insightful and writes in a clear compelling way. GPT-5.6-Sol on the other hand is extremely diligent, I can give it a list of 8 things to do and you will be sure that they will be done.
Fable feels more arrogant to me, I was both to get it to build a new benchmark for me - 5.6 worked between 6 hours and 2 days (I tried several times) and it came up with very thoroughly tested, working benchmark. Fable came back within 40 minutes (twice) and the benchmark sounded smart, but was ultimately was 'vibe' based slop and since it was Fable's vibes that was doing the judging, it decided that it was good to go (it kept giving Fable 100% score btw).
Some thoughts by category:
UI & App building: Fable will still craft a better UI from scratch, the flow of the app would probably be a bit nicer. But I find that Fable often misses quite key things, which GPT-5.6-Sol doesn't. GPT's Frontend skills are big jump vs previous GPT models, but still not as great overall.
Writing: Fable is better hands down, Sol feels quite difficult to align to what I want to say or explain things to me simply. Though I think the 'Pro' model writes clearer.
Robustness & Reliability: This is where I think GPT-5.6-Sol wins for me hands down. Fable seems to do things of high quality, but I can never relax with it, it always misses something. With 5.6 this just almost never happens.
Other things where I liked GPT-5.6-Sol, but can't compare to Fable directly.
- Video editing is actually working now, it is not completely perfect, but with the right skill/guidance you can just give it 1h footage and it can give you a 5 min highlight clip no problem
- Computer use - getting really rather good, very usable
- Sub agents - it is very fluent at managing sub-agents and speaking to different threads, can help with some new workflows
- Adhering to existing code patterns - I love this, even without asking it would implement something in a way that aligns with you app - major problem for slop generation
- Research - I think it is getting quite a bit better, it still has some bad patterns (e.g being too tactical), but it feels like it is more steerable to be a good researcher
- Multi-day runs - the /goal feature is pretty insane with 5.6-Sol, you can run it for days if you wanted to and it does work. Useful to have another thread or /side to check up on it, but I have some great results with it
- Token efficiency - it is so much more token efficient and faster than 5.5, in reality it is now much faster than Fable too
On the downside, you can feel that Fable is naturally smarter, and I did have some baffling moments with 5.6 when I was getting it to make a fairly simple change in 8 turns - it seemed to get stuck in a dumb stream that was hard to get out of. So it is not AGI, don't get too carried away by the hype.
I have some phenomenal examples that I'm honestly blown away by that I'll share, but as a side anecdote, I have a kind of 'swear meter' which counts how often I'm rude to Codex. In GPT-5.5 era, the % was at around 4-5%, it dropped to 1-2% when I was testing GPT-5.6-Sol and it shot up to 7% when I went back to 5.5 - it was so shocking to go back to 5.5 and experience how much worse it was.
So is GPT-5.6-Sol better than Fable? On pure intelligence - no. But man, I missed it when I just wanted to get sh*t done. It is insanely capable workhorse that you can give any task to and just expect it to be done. No lectures or 'you are absolutely rightisms', nothing is beneath it, if it takes 2 days to do some dirty work, it will do it.
It feels like the first time in a while when we have quite different types of frontier intelligences that benchmark sort of similarly, but feel very different. If you can, you would be probably better off using both and iteratively finding what you'd use Fable or GPT-5.6-Sol for. Perhaps, something like - an architectural discussion with Fable, implementation with 5.6 and docs & comms with Fable.
Show more
I spent a LOT of time through the hardest 3D prompts at Fable, it is a 45 min video, but I have 60+ very cool demos for you. Also prompts in the next post.