Register and share your invite link to earn from video plays and referrals.

Search results for 551蓬莱〉の
551蓬莱〉の community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 551蓬莱〉の
Fable 5.1 is under-hyped because it's so expensive It is VASTLY SUPERIOR to Opus 5.5 on very complex codebases Not just that, it one shots hard bugs fixes and features without a lot of prompt itreation Opus 5.5 shines in contrast to Opus 5 which was a terrible and insanely expensive model
Show more
opus 5.5, fable 5.1, and gpt-6 astra — and we're only comparing the public models here both openai and anthropic have more advanced internal models. oai showed internal-model performance graphs around the navier-stokes work, which could give us a rough idea of how far ahead they are internally probably around 2 months ahead?
Show more
Anthropic’s new Opus 5.5 beats Fable 5.1 on AAII at ~2.5x lower cost. An hour later OpenAI releases GPT-6 Sol and Luna, both at half price of 5.6 predecessors. This is an unprecedented jump in performance and drop in price for frontier intelligence.
Show more
opus 5.5 is an RL run of the 5 pre-train i think even anthropic was surprised by how good it turned out. apparently it was first going to be called 5.1, then 5.2 as the results kept getting better, and eventually 5.5 because it ended up much stronger than expected it took around 2 months to get there
Show more
Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited? It's a combination of two things: Opus's efficiency, and the 50% limit on Fable. Claude Code subs are very generous with their usage, but only half is allowed to be used by Fable. Separately, Opus is way more gentle with costs. Opus 5.5 on High is over 2x cheaper than Fable 5.1 High. The result is that, roughly, going from Fable 5.1 high to Opus 5.5 high is a 4.3x increase in limits. Fable 5.1 xhigh to Opus 5.5 high (move I made) is a 6.6x increase 🤯
Show more
0
110
1.3K
30
Forward to community
Opus 5.5 genuinely feels like the best model release from any frontier lab so far. Less slop and way more human feeling outputs particularly for technology and stock research. I'm actually enjoying reading its outputs compared to Astra and even Fable 5.1 which dumps out stuff that reads like a high school textbook. I'm very impressed. This is ultra bullish for Anthropic.
Show more
ANTHROPIC: CLAUDE OPUS 5.5 PERFORMS AT LEVEL OF CLAUDE FABLE 5.1 ON MOST WORK, COSTS AROUND 40% LESS TO RUN THAN OPUS 5
🚨 OPUS 5.5 Officially Drops today > Opus 5.5 outperforms Fable 5.1 > And around 20% cheaper than Opus 5 and Fable 5.1 > Could finally make Claude's flagship models much more affordable > Also it's better than current Gpt-6 Astra We are so back And also Sol is coming Big night
Show more
🚨Opus 5.5 output is TERRIFYING This is actually insane Opus 5 is apparently being ROUTED to Opus 5.5 for some users Opus 5.5 is making GPT-6 Astra and Fable 5.1 look way less impressive What the hell is Anthropic cooking?
Show more
SWE-Together Update: Claude Fable 5 and 5.1 are the new top models on SWE-Together, and Meta's Muse Spark 1.3 is the best value on the board. SWE-Together is our benchmark of 109 real coding tasks with a simulated user in the loop, built to tell you whether a new model will actually match your expectations before you hand it your work. Six quick findings from the updated board: 1. Fable 5 beats 5.1 on consistency, but 5.1 is more efficient. At their best the two actually perform the same. The gap is in the worst runs, where Fable 5.1 scored zero on 14 trials and Fable 5 on only 6. However, 5.1 is 25% faster, about 2x cheaper, and needs slightly fewer corrections from the user (1.45 vs 1.53 corrective messages per task) than Fable 5. 2. Muse Spark 1.3 is the value outlier. Per solved task, it is about 2.5x cheaper than Fable 5.1 and 5x cheaper than Fable 5. Performance-wise, it ties with Fable 5 and 5.1 for best on Python (73%) and Rust (100%). 3. Newer is not automatically better at collaborative coding. GPT-5.6-Sol is the newer model, yet it needs more corrections from the user than GPT-5.5 (1.66 vs 1.59 per task). 4. "Stronger models need less steering" is a trend, but not guaranteed. With more models added, the correlation between pass@1 and user correction weakens from -0.92 to -0.72, and Fable 5 takes more corrections than Opus 4.8 (1.53 vs 1.38) despite a 6 pp higher pass@1. 5. Frontier progress contributes greatly to stability. Fable 5 converts 89% of its pass@1 into pass² (both runs solve), versus 60 to 67% for DeepSeek V4 Pro, GLM-5.1, and MiniMax 2.7. 6. There is still plenty of headroom. 16 of 109 tasks have zero solves across all tested models, and even Fable 5 at 69% pass@1 sits about 9 pp below the ~78% that the original human patches scored. The full leaderboard, with per-task and per-trial breakdowns, is at and the benchmark is open source at
Show more