Register and share your invite link to earn from video plays and referrals.

follistatin
@follistatindev
25 | swe ml | hobbies: game dev, bodybuilding, powerlifting, running, biohacking, a friend to ai | 840 SQ | 550 B | 655 DL | 440 OHP
1.5K Following    259 Followers
Is there any model truthfully cheaper than Deepseek V4.1-Flash? its honestly cheaper than even the 27b models...
>sonnet >$7.6 per task to run the most expensive model on the entire leaderboard what are we even doing atp
Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we’ve seen With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and to #2# on the Intelligence Index behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens, however it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5’s Cost per Task) Key takeaways: ➤ Meets leading models on agentic terminal use and knowledge work: in Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it ➤ Heaviest token use we have measured: at max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured on around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max) ➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol. At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task ➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: as a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon. Other model details: ➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5 ➤ Pricing: unchanged from Sonnet 5’s latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2 ➤ Effort settings: five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in TerminalBench 4.0, falling back to Sonnet 5 in all cases.
Show more
GPT-6 Luna is awful. Even a local 9B model could do better.
If you had asked me a year ago I would have predicted that an LLM solving a millennium problem would shut these guys up but it hasn't even slowed them down a little bit. Never underestimate the power of exponential growth or the depths of human stupidity.
Show more
M3.1-Flash is the worst performer, and gpt-6-astra is astonishingly bad at this, only M3.1-Flash was worse Let me know in DMs if you want me to test more models on this private gamefeel/gamedev benchmark. Only tested one subsection of Opus 5.5 on this as I ran out of credits.
Show more
Introducing: The gamefeel™ benchmark a private benchmark made by me with dozens of gamedev tasks: focusing on game balancing and game feel, and understanding of game mechanics (roguelikes, MMORPGs, strategy games, survival/sandbox games included) #anthropic# #openai# #llm#
Show more
"It's not our war" well they just hit one of our embassies so I think it's starting to be our fucking war
0
540
204.4K
7.2K
Forward to community
I used to debug code now I type this and hit enter
0
179
23.5K
1.1K
Forward to community
yeah opus 5.5 is just it "it just works" they fixed slopus vocab, is now much less likely to over abstract, digs deeper when needed only sol 6 feels off, have to constantly poke, its a slop MONSTER it tries to write a for loop to make an edit wtf man ant im coming home
Show more