Register and share your invite link to earn from video plays and referrals.

forloop
@forloopcodes
tokenmaxxing inferencel looping agents, or perhaps a promptchud. give back my employment email: forloop@poke.com contextplus (2k), cargo install safeinstall
Joined September 2024
2.1K Following    9.6K Followers
GPT 6 Astra is estimated to have ~1.2T Active Parameters and ~10T Total Parameters. Reports indicate pre-training for GPT-5.5 (codename Spud) concluded on March 24 at the Abilene campus. Astra's pretraining run (Doug) would have taken over that cluster after that, OpenAI confirmed the math results were achieved by an internal version of Astra on August 1. So pretraining ran roughly late March to late June/early July, that is ~90–110 days. Abilene's operational figure was approximately 150,000 to 200,000 GPUs as of May 2026, with inference sharing the site. So ">100k" for the run means 100 to 150k GPUs. Around August 25, information circulated that OpenAI had completed a base model called "Bel" with over 10 trillion parameters, successor to Doug. Bel finished ~Aug 25 but Astra was already doing math Aug 1, so Astra is almost certainly on Doug, and Bel is a later model. Compute for pretraining alone: 120k GPUs × 2.25e15 (BF16) × 0.30 MFU × 100 days ≈ 7e26 FLOP. Range 4e26 to 1.5e27 depending on FP8 and MFU. Whole 100 days on pretraining, which is now plausible since RL ran afterward (the big RL run was restarted in August per Vellum/CNBC reporting). Now the data wall does real work. Solve N_active = C / (6·D): Tokens D N_active at 7e26 40T 2.9T 60T 1.9T 80T 1.5T 100T 1.2T Realistic 2026 corpora with synthetic data and repeats top out ~60–100T, so active params land at 1 to 2T. Even at the conservative 4e26 end with 80T tokens, you get ~830B active. Astra bills $50/M output at 62 t/s vs Sol's $30/M, a 1.7x price step, consistent with roughly 1.5 to 2.5x active params over its predecessor, not a 10x jump. Rough, but it argues against the >2T end. Active parameters: ~1T–1.5T (central ~1.2T) Total parameters: ~10–14T assuming 8–12x MoE sparsity (GPT-4's leaked ratio was ~6.5x; modern large MoE runs 10x+)
Show more
@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.
Show more
0
58
2.1K
104
Forward to community