Register and share your invite link to earn from video plays and referrals.

Louis-François Bouchard 🎥🤖
@Whats_AI
Training AI Engineers on YouTube, Substack and our courses. Co-founder @towards_ai. Ex-Ph.D. student @Mila_Quebec. On the road to 100k on YT this year 👇
6.2K Following    11.4K Followers
5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug. Claude Opus 5.5 is the new #1# on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started running this in June 2026. It is also the first model to clear 91 out of 100 on our rubrics. One thing to note: at max effort it takes 17 minutes and $3.43 to write one script. It is the slowest configuration on the entire board. By FAR. Not the most expensive one, but in the top 5 most expensive. More details on the effort and thinking levels of Opus 5.5 👇
Show more
0
67
1.4K
67
Forward to community
We tested 146 model variants on 10 script-writing tasks, 5 runs each, scored by a 3-judge panel on tone, voice, and whether a human would actually like it enough to read it out loud without cringing. Our goal is to measure creative writing quality. Not agentic or coding performance, or even one-shot generation performance. We provide a very detailed prompt with all the info the model needs to do its job (in theory). Here are some insights from the main family of models with GPT, Fable, Gemini, Muse... (see image for more info): 1. If you are using sol 5.6, skip the ultra setting. It doubles the bill from $0.21 to $0.43 per task and buys only 66 Elo. 2. Extra thinking is worth paying for on the small models, not the big ones. Luna gains 230 Elo for 3× the cost. Sol gains 66 for 2×. Grok 4.6 and Gemini 3.1 Pro get worse when you turn it up. 3. Claude Opus 5 is the value pick of the whole chart. Max effort is $0.42 a task, and it out-writes every GPT config on the board — including GPT-6 Astra, which costs 8× more. I know this one is controversial, but I also see it using it personally. And that's what you should trust in any case: you own usage, not a benchmark. 4. Never pay for Sonnet 5's max effort. It burns $2.96 a task, 7× what Opus 5 max costs, and writes 216 Elo worse. 5. GPT-6 Astra is the most expensive way to be mid. $3.42 a task at max effort, 246 Elo behind Opus 5 max. 6. Meta Muse Spark 1.3 is the budget outlier: $0.034 a task, and it out-writes every Gemini 3.1 Pro config — which costs 4× more. 7. Claude Fable 5 still tops the board. $4.36 a task at max effort: 10× Opus 5 max for 72 Elo. Pay it only if the words are the product or you still have your subscription tokens. 8. Claude Fable 5.1 is not a writing upgrade. At max effort it costs 25% more than Fable 5 ($3.15 vs $2.52) and scores 6 Elo lower, which is inside the noise band.
Show more
This 1 hour workshop on context engineering will help you build agents that hold long conversations and cut your token cost by up to 90%. Last month we gave it live at the AI Engineer World's Fair 2026 in SF. Now it is free on their YouTube channel. And everything is open-source. Context engineering is deciding what your model sees every time you call it. What stays in the context window, what gets dropped, and when. Get it wrong and your agent forgets things, answers slower, and costs more with every message. My colleagues @omar_solano1, @samridhivaid and I from @towards_AI taught all of it using our AI Tutor App, the product our students ask questions in while going through our courses. Here is what we covered inside the workshop: → 11 ways to manage chat history, and which one won → Why summarizing the conversation often costs more than keeping all of it → How prompt caching cuts around 90% of your token cost → How we cut our bill by moving to a cheaper model → What broke when we tried running it on a laptop This is a complete masterclass on context engineering and after watching it you will know how to handle it in your own agent. Thanks @swyx and the @aiDotEngineer team for having us (once again) 🙏 Workshop video:
Show more
Big news from our internal writing benchmark (early results): Claude Opus 5 by @AnthropicAI is now #1# for writing in our editorial voice, at 2817 Elo, surpassing Claude Fable 5 and Kimi. Already! That is a jump from #15# to #1# over its predecessor (if we take all thinking variants into account), Opus 4.8, at the same API price. Reasoning effort actually matters this time. At default effort it lands #6#. At max effort it takes the top spot, thinking for over three minutes per script. Seems obvious, but it wasn’t the case for 4.8, though it is for Fable.
Show more
Big news from our internal writing benchmark (early results): Kimi K3 by @Kimi_Moonshot is now #1# for writing in our editorial voice, at 2840 Elo, surpassing Claude Fable 5. That is a jump from #21# to #1# over its predecessor, Kimi K2.6. And it runs at about $0.25 per script: 5x cheaper than the model it just displaced at the top. First time an open-weights model tops our board. It surprised us too. More in the replies 👇
Show more