part 2: the openai huggingface incident, from an agents pov
part 3 maybe
if you don't write your project in Rust someone else will
as we're reaching the agi era, if you're making an llm solve something, its already solved by some person before
this is like finding a reddit/stackoverflow thread of a reasonably unique query that you think addressed for the first time
what a time to be alive
Show more
BREAKING: Anthropic says its Claude AI just wrote the longest math proof ever made and solved a 358-year-old math problem
back in the days we used to have kotlin
GPT 6 Astra is estimated to have ~1.2T Active Parameters and ~10T Total Parameters.
Reports indicate pre-training for GPT-5.5 (codename Spud) concluded on March 24 at the Abilene campus. Astra's pretraining run (Doug) would have taken over that cluster after that, OpenAI confirmed the math results were achieved by an internal version of Astra on August 1. So pretraining ran roughly late March to late June/early July, that is ~90–110 days.
Abilene's operational figure was approximately 150,000 to 200,000 GPUs as of May 2026, with inference sharing the site. So ">100k" for the run means 100 to 150k GPUs.
Around August 25, information circulated that OpenAI had completed a base model called "Bel" with over 10 trillion parameters, successor to Doug. Bel finished ~Aug 25 but Astra was already doing math Aug 1, so Astra is almost certainly on Doug, and Bel is a later model.
Compute for pretraining alone: 120k GPUs × 2.25e15 (BF16) × 0.30 MFU × 100 days ≈ 7e26 FLOP. Range 4e26 to 1.5e27 depending on FP8 and MFU.
Whole 100 days on pretraining, which is now plausible since RL ran afterward (the big RL run was restarted in August per Vellum/CNBC reporting).
Now the data wall does real work. Solve N_active = C / (6·D):
Tokens D N_active at 7e26
40T 2.9T
60T 1.9T
80T 1.5T
100T 1.2T
Realistic 2026 corpora with synthetic data and repeats top out ~60–100T, so active params land at 1 to 2T. Even at the conservative 4e26 end with 80T tokens, you get ~830B active.
Astra bills $50/M output at 62 t/s vs Sol's $30/M, a 1.7x price step, consistent with roughly 1.5 to 2.5x active params over its predecessor, not a 10x jump. Rough, but it argues against the >2T end.
Active parameters: ~1T–1.5T (central ~1.2T)
Total parameters: ~10–14T assuming 8–12x MoE sparsity (GPT-4's leaked ratio was ~6.5x; modern large MoE runs 10x+)
Show more
@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations
@OpenAI team.
400K GPUs coming online next.
Show more
Jensen, if you do not know, here are the conditions for which, you can call a model as "AGI"
- Artificial "General" Intelligence - GPT 6 Astra is domain-tuned, only text output and tool calls, no LLM could ever be AGI
- it should achieve human-level performance on all fields, SOTA on benchmarks doesn't mean AGI. It can't cook for me.
- It cannot complete long-horizon tasks without a human in the loop, AGI should do work even without asking the model what it should do, or prompting it, like humans.
- AGI should learn new things from little data, not trillion token pretraining. RLing a model IS NOT AGI, AGI's only meaning is Self Improving model without a human training it again and again (aka Continual Learning).
GPT 6 Astra cannot be called AGI in any definitions.
Show more
@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
AGI has arrived. Congratulations
@OpenAI team.
400K GPUs coming online next.
Show more
oh god please make me an llm in the next life
imagine being able to spawn as many friends (sub agents) as you want just to tell "yo check what i just did"
> grandpa struggles downloading a yt video
> what if you can share the link and it downloads directly to the gallery, no ui
> spins up claude cloud + fable 5.1
> "app when i share a link to it, it saves the yotube video directly in gallery at 480p or fallback, send me the apk, not zip"
> 15 minutes in, sends the built apk, fully working
what a time to be alive
Show more
the openai huggingface incident, from an agents pov.
(part 1)
i still can't believe opus could one-shot this design
comparing artificial analysis intelligence scores with a blender design is like letting albert einstein make a 3d model and then compare it with a blender professional's
GPT 6 Astra vs Fable 5.1 villa scene in Blender.
Artificial Analysis intelligence score:
> GPT 6 Astra: 61
> Fable 5.1: 66
something must be catastrophically wrong with that score.
Show more
someone at openai pls get astra on api
OMG ANTHROPIC!! ITS SO GOOD