Register and share your invite link to earn from video plays and referrals.

Enactra AI
@EnactraAI
Agentic world simulation for physical RSI
575 Following    377 Followers
Insanely good showing for DeepSeek it's still not a great vision model, but clearly Visual Primitives help a lot here. I wish they re-released the paper…
SWE‑2 reshapes BuildingBench’s cost–quality frontier 👀🏗️ @cognition's SWE‑2 (max) scores 67.3—just 2.5 points behind GPT‑6 Sol (ultra, 69.8), and ahead of GPT‑5.6 Luna (61.6). Free to use for now, it reshapes the budget end of the cost–quality curve: building in 3D just got more accessible. It's really a good model. Congrats @cognition @ScottWu46 @jaredpalmer ! How much can you build on a $0 model budget? 🚀
Show more
A lot of the work behind DeliveryGym is systems side: making long-horizon agentic online RL practical in UE5 while balancing photorealistic rendering with rollout throughput. We use SPEAR for the Python interface to Unreal, alongside our distributed rollout infra: parallel simulator workers, asynchronous coordination across environment execution, inference, and policy training. Thanks @mikeroberts3000 and the SPEAR team for building and sharing these tools!
Show more
RT @koe_ye40329: Excited to share that our SimWorld-Studio paper has been accepted as a poster at #NeurIPS2026#! 🎉 Congrats and thanks to o…
From BuildingBench to a square kilometre of NYC 🗽 Opus 5.5 vs. Astra, reconstructing the area around Madison Square Park. Given the same maps, terrain, and photographs, each agent rebuilt the scene independently in Blender and Unreal Engine. No internet. No human help. Buildings, pedestrians, and traffic—though Opus’s pedestrians glide through the park without moving their feet. Why walk when you can float? 😂 Opus leads on BuildingBench. Which model builds the more convincing city? 🏙️👀
Show more
New leader alert 🚨 Opus 5.5 takes the BuildingBench crown, while GPT‑6 Astra and Sol stand out on cost efficiency. 🏗️ 🥇 Opus 5.5 (max) jumps from Opus 5’s 74.5 to 86.8—the highest score yet, ahead of GPT‑6 Astra (ultra, 84.3) and Fable 5.1 (max, 81.4). 💸 GPT‑6 Astra stays close at a fraction of the cost: $7.05 per building vs. ~$45.47 for Opus 5.5. Just 2.5 points behind, at 84% lower cost. ⚡ GPT‑6 Sol takes a different tradeoff: • Ultra: 69.8 at ~$1.92/building • Max: 69.6 at ~$1.89/building Below GPT‑5.6 Sol’s 73.9, but roughly 75% cheaper than its $7.54/building. Both settings join the benchmark’s cost–quality frontier. Watch the leaderboard shift 👇 Building comparisons from Opus 5.5, GPT‑6 Astra, and GPT‑6 Sol in the thread!
Show more
What do those scores look like in 3D? 👀🏗️ Opus 5.5 (max) vs. GPT‑6 Sol (ultra) vs. GPT‑6 Astra (ultra), side by side. Median cost per building: 🟣 Opus 5.5: ~$45.47 🟢 Astra: $7.05 🔵 Sol: ~$1.92
Show more
Opus 5.5 raises the bar for quality. GPT‑6 Sol lowers the cost of building in 3D. 🏗️👀
What do those scores look like in 3D? 👀🏗️ Opus 5.5 (max) vs. GPT‑6 Sol (ultra) vs. GPT‑6 Astra (ultra), side by side. Median cost per building: 🟣 Opus 5.5: ~$45.47 🟢 Astra: $7.05 🔵 Sol: ~$1.92
Show more
New leader alert 🚨 Opus 5.5 takes the BuildingBench crown, while GPT‑6 Astra and Sol stand out on cost efficiency. 🏗️ 🥇 Opus 5.5 (max) jumps from Opus 5’s 74.5 to 86.8—the highest score yet, ahead of GPT‑6 Astra (ultra, 84.3) and Fable 5.1 (max, 81.4). 💸 GPT‑6 Astra stays close at a fraction of the cost: $7.05 per building vs. ~$45.47 for Opus 5.5. Just 2.5 points behind, at 84% lower cost. ⚡ GPT‑6 Sol takes a different tradeoff: • Ultra: 69.8 at ~$1.92/building • Max: 69.6 at ~$1.89/building Below GPT‑5.6 Sol’s 73.9, but roughly 75% cheaper than its $7.54/building. Both settings join the benchmark’s cost–quality frontier. Watch the leaderboard shift 👇 Building comparisons from Opus 5.5, GPT‑6 Astra, and GPT‑6 Sol in the thread!
Show more
See the leap from Grok 4.6 to 4.7 for yourself🧐 Missing walls, inside-out surfaces, blurry textures… Grok 4.7 fixes much of this, producing more faithful, complete, and detailed buildings. On BuildingBench, both at xhigh: ➤ Score: 0.696 → 0.783 (+12.5%) ➤ Median cost per building: $10.59 → $11.43 (+8%) A substantial quality jump for a modest cost increase. It's surely a strong release! @milichab @veggie_eric @elonmusk
Show more
Grok’s leap isn’t just in coding—it’s showing up in 3D simulation, too 🏗️ Grok 4.7 (xhigh) scores 0.783 on BuildingBench, taking Grok from #9# to #3#! A big jump from 4.6’s 0.695. It now trails only GPT‑6 Astra (ultra, 0.843) and Fable 5.1 (max, 0.814). Median cost per building: $11.43, versus Astra’s $7.11 and Fable 5.1’s $33.15. Around 66% cheaper than Fable. Congrats @SpaceXAI and @ElonMusk! Coding agents are getting better at building the world.
Show more
@DarthJML Building bench seems to be a benchmark more down that alley.
Grok’s leap isn’t just in coding—it’s showing up in 3D simulation, too 🏗️ Grok 4.7 (xhigh) scores 0.783 on BuildingBench, taking Grok from #9# to #3#! A big jump from 4.6’s 0.695. It now trails only GPT‑6 Astra (ultra, 0.843) and Fable 5.1 (max, 0.814). Median cost per building: $11.43, versus Astra’s $7.11 and Fable 5.1’s $33.15. Around 66% cheaper than Fable. Congrats @SpaceXAI and @ElonMusk! Coding agents are getting better at building the world.
Show more
New episode of Running Man 🏃 DeepSeek V4.1 Flash vs. Union Alpha—three real-time rounds through simulated NYC. Final score: 3–0. No spoilers—watch to see who gets swept 😂 Which model should join the battle next? 👉
Show more
Astra vs. Fable 5.1: tag, you’re it! 🏃🗽 Tag Game, all in real time, in the NYC simulation Astra itself created. The catch? The world doesn’t pause while they think. Cue some awkward collisions with pedestrians and cars 😂 Frontier agents are stepping into embodied worlds, and a simple game of tag puts spatial awareness, planning, opponent prediction, and timely reactions to the test. Skills they’ll need to thrive beyond the chat or coding window. Watch Fable 5.1 beat Astra 2-1 👑👀
Show more
Fun to see coding agents being deployed in physical-world simulations for real-time tasks. The physical RSI loop is starting to run, literally. 🏃🔁
Finally—and seemingly all at once—coding agents are breaking out of the chatbox and into real-time worlds. Can’t wait to see what they’ll achieve in a year, or maybe just a month!
Astra vs. Fable 5.1: tag, you’re it! 🏃🗽 Tag Game, all in real time, in the NYC simulation Astra itself created. The catch? The world doesn’t pause while they think. Cue some awkward collisions with pedestrians and cars 😂 Frontier agents are stepping into embodied worlds, and a simple game of tag puts spatial awareness, planning, opponent prediction, and timely reactions to the test. Skills they’ll need to thrive beyond the chat or coding window. Watch Fable 5.1 beat Astra 2-1 👑👀
Show more
Union Alpha (stealth model): same building performance as DeepSeek 4.1 Flash, with only 1/10 the tokens. 🤯👀 Watch it alongside DeepSeek, GPT‑6 Astra, and Kimi K3 below. Free to try for a week on OpenRouter, OpenCode, and Cline. Who’s behind this one? 🕵️
Show more
Coding agents like GPT‑6 Astra and DS V4.1 Flash are entering the 3D world. We need new benchmarks for their spatial reasoning. 🏗️ Introducing _BuildingBench_: can coding agents turn real-world images into coherent 3D buildings, and at what cost? BuildingBench is a step toward coding agents creating reusable worlds that people and other agents can inspect, edit, and build on for design, games, and physical simulation. And the quality–cost frontier is moving fast: 👀 🔥 GPT‑6 Astra achieves the highest score in our comparison at ~80% lower cost than Fable 5.1. 🔥 DeepSeek V4.1 Flash delivers strong quality at a median model cost of just $2.31 per building. Watch the frontier shift. More details, interactive leaderboard, and GitHub in the thread 👇
Show more
Explore BuildingBench 🏗️: 📖 Blog: 🏆 Interactive leaderboard: 💻 GitHub: Inspect the buildings behind the scores. 🕹️
Show more
Can Coding Agent Reconstruct the World, Building by Building?
Explore BuildingBench 🏗️: 📖 Blog: 🏆 Interactive leaderboard: 💻 GitHub: Inspect the buildings behind the scores. 🕹️
Show more