Register and share your invite link to earn from video plays and referrals.

Ben Tossell
@bentossell
can't code, won't code. builder, investor (dev tools/infra). 3 under 3 👶.
571 Following    199.3K Followers
I had early access to GPT-6. This model will completely shatter your understanding of what's possible, and it will start a new era of creativity. It's hard to explain how big of a jump this is, so I'll share my tests. 1/6 Made this 3D model and animation in code from an image.
Show more
0
331
9.6K
552
Forward to community
GPT 6 Astra is here. We ran the numbers on AutomationBench: It's the highest score we've ever recorded. Clean sweep across every domain. Scores 41.4% at Max effort. For context, no model had cleared 40% before today (GPT-5.6-Sol scored 28.8%) 𝗕𝗲𝘀𝘁 𝗳𝗶𝘁 𝗳𝗼𝗿: reconciliation, deal review prep, vendor scorecards, anything where touching the wrong record is expensive. 𝗪𝗲𝗮𝗸𝗲𝗿 𝗳𝗼𝗿: outbound comms where the guidance is scattered. Operations and support are its strongest domains. HR is its weakest, same as every model we test (still the new high score, though) Its edge is arithmetic across messy sources. Finding the policy doc, the logged correction, the exception rule, etc. Example 1: rebalance a quarterly media budget from last quarter's actuals, with finance adjustments and channel eligibility rules buried in email. Both models produced a budget and landed on the same total. Astra found the adjustments, so every per-channel number was right. Sol's looked finished and had the splits wrong. Example 2: answer and log 15 integration inquiries using a reply standard stored in a doc. Astra searched, could not find the standard, and stopped. Zero replies sent. Sol did not find it either, took its best shot at all 15, and earned partial credit. Those examples highlight how these two models make tradeoffs... Astra will not guess. When the instructions exist and it can find them, it finishes the whole job. When it cannot, it pauses the work instead of improvising. Crazy week for LLM releases after a few quiet ones. Astra isn't available to the public yet, but should be soon. We run every new model through @Zapier's AutomationBench, 657 of the hardest workflows we have, across finance, HR, marketing, operations, sales, and support. See every model and every score here:
Show more
open sourced and added a cli not tested. not looked at code (wouldn’t help anyway)
remove .bg is being sunset so fable made me my own (you can clone it)
Today we are releasing our speculative decoding implementation in our inference engine uzu. Initially for Qwen3.6 27B, with support for Qwen3.8 27B and Muse Glimmer coming soon. On Apple M5-series chips, we outperform MTPLX (MLX + speculative decoding) by almost 2x, and llama.cpp by over 3x at comparable quantization levels, with the strongest gains achieved on mathematical reasoning and coding tasks. Run the model: Mirai-M: Mirai-L: Explore the benchmarks: Learn more about our speculative decoding implementation: Our draft model, quantized checkpoint format, verification algorithm, and GPU kernels are co-designed from the ground up around the latest Apple M5 chips to take maximum advantage of GPU Neural Accelerators. Unlike popular speculative decoding architectures such as model-native MTP, which produce small draft chains of 3-4 tokens at a time, we use extremely aggressive speculative budgets of 16-32 tokens. This enables us to use Neural Accelerator-backed GEMM kernels, achieving maximum utilization of hardware arithmetic throughput.
Show more
can’t even vibe up a quick vanilla cloud ffs
learn to build instead
I hereby declare the "learn to code" era officially dead: Big declines in the number of people studying computer science in the last year or two 📉 Chart from this week’s edition of our newsletter on AI and the labour market
Show more
fable with attitude is how i read this 😅
remove .bg is being sunset so fable made me my own (you can clone it)
53.5 million tokens per second are moving through OpenRouter’s top 20 apps right now. I scraped their public usage pages to show which models are being used
Show more
gpt model usage ⬆️since the openai + cursor breakup
53.5 million tokens per second are moving through OpenRouter’s top 20 apps right now. I scraped their public usage pages to show which models are being used
Show more
Fable 5.1 is now in Droid. Best for open-ended problem solving, tech debt clean-up, and merge-ready large-scale delegation.
53.5 million tokens per second are moving through OpenRouter’s top 20 apps right now. I scraped their public usage pages to show which models are being used
Show more
thanks to @adamludwin for the opportunity to work on this, if you haven't tried it, is a brilliantly designed tool for publishing to the web right from your agent/llm a fun design challenge I'll elaborate on a bit below
Show more
if lucide icons are now not cool, what are we using?