Register and share your invite link to earn from video plays and referrals.

Lisan al Gaib
@scaling01
lead them to paradise LisanBench: Impressum & Datenschutz:
1.3K Following    60.1K Followers
OMG SAVED Opus 5.5 is NOT looping
show me the no-CoT time-horizons you god damn loopers
as i said, people are sleeping on bytedance
The Chinese AI Infrastructure Boom: Introducing the SemiAnalysis China Datacenter Model 1,000+ facilities across 60+ operators mapped, built retail-first and flipped by AI, largest hyperscaler leases 1/5 national capacity, 100MW in 12 months, Eastern Data Western Compute
Show more
Claude Opus 5.5 (xhigh) scores 31.2% on WeirdML v3, a clear advancement over Fable 5.1 at 26.0%, but well behind GPT 6 Astra at 42.2% It also comes out at less than half the price of Opus 5.
Show more
I just talked to one of my friends who was super anti AI but after trying Opus 5.5 he's now AI pilled and likes vibe coding
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62% Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8). As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts. Key takeaways: ➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom ➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks ➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
Show more
Claude Opus 5.5 has now the highest score on SimpleBench with 88.4%
Opus 5.5 makes me think Astra is even smaller than I thought like it's insane how good Opus 5.5 is
Claude Opus 5.5 High vs GPT 6 Astra Max > 640×360 pixel art of cheetah running look at the difference
Opus 5.5 is the best "vision" model Anthropic ever made better than Fable 5, better than GPT-6 Sol, worse than GPT-6 Astra vs Fable 5.1 (high): - estimated cost: ↓ 60% - detection: ↑ 11.84 pp - reasoning: ↑ 12.80 pp vs GPT-6 Sol (high): - extraction: ↑ 11.00 pp - reasoning: ↑ 7.95 pp ↓ 16 more examples
Show more
it's called the forbidden axis for a reason well, except it's no longer forbidden. OpenAI opened the floodgates and will release their second looped language model besides GPT-6 Astra on DevDay
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
Show more
probably my favorite quote of all time
it's called the forbidden axis for a reason well, except it's no longer forbidden. OpenAI opened the floodgates and will release their second looped language model besides GPT-6 Astra on DevDay
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
Show more
Just dropped my new article. Hope you enjoy it :)
mind you there are no children watching any of this this presentation is for grown adults lmao
Alexander Wang (Meta): "pretty soon we are dropping the most capable model we have ever trained" why not today? :(
Alexander Wang (Meta): "pretty soon we are dropping the most capable model we have ever trained" why not today? :(
We’d like to know how far current frontier AIs generalize “out-of-distribution” (that is, how able they are to solve problems they haven’t seen before), to know how fast things will move. So we watched some AIs play Pokemon.
Show more
looks like the big chungus muse models are going to be announced today
Just dropped my new article. Hope you enjoy it :)
I literally can't talk with Opus 5.5 about my article lmao the safety monitors just go off I guess I wrote an infohazard