Register and share your invite link to earn from video plays and referrals.

Search results for VUCCA
VUCCA community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including VUCCA
Vulcan Wet Dress Rehearsal (WDR)✔️ Yesterday, ULA conducted a Wet Dress Rehearsal (WDR), at Space Launch Complex-41. This full day-of-launch demonstration loaded propellant into the Vulcan rocket and new LEO-optimized Centaur V upper stage, and tested functions of the new Vulcan Launch Platform (VLP)-A, the second Vulcan platform to serve @AmazonLeo missions on our heavy-lift rocket.
Show more
Josephine Vaccarello Named President of MSG Entertainment
The journey building VulcanBench Safety v1 begins, and with it, my first private repo VulcanConduct. This is phase zero of my AI safety eval suite build out, more to come, will share as I go as always 🖖
Show more
Chugging along with @VulcanBench Cyber Eval Suite v1...just crossed 53,000 loc, and quite a bit more to go.
Grok 4.5 still ranks #1# on VulcanBench Eval Suite 3.....a benchmark built specifically around hard, real-world software engineering work It tests models on 23 frontier-hard engineering tasks taken directly from real merged open-source PRs And Grok 4.5 is still sitting at the very top.....ranked #1# Grok’s engineering capability is insanely strong
Show more
Building new language-specific eval suites for @VulcanBench and making good progress on my Python eval suite v1. And yeah, this one is going to be a monster 🐍
Launching my third attempt at my @VulcanBench model x harness benchmark on eval suite 3. I think I finally have all the guardrails in so cheating will be guaranteed to be impossible. Not going by the honor system, I want to know for sure there no cheating. Oh, and kicking this off remotely on my M1 Mac Studio from my phone, using Claude Code, from the lake, in between paddle boarding sessions 🏄‍♂️ 🤘
Show more
Grok 4.5 just ranked #1# on VulcanBench, outperforming Claude Fable 5, GPT-5.6 Sol and Kimi K3 across 23 real-world merged PRs in Python, Rust, TypeScript, JavaScript and Go Insane efficiency: • 21/23 tasks solved - 91.3% • Best result across all 11 configurations tested • Just $0.32 per solved task Claude Fable 5 and GPT-5.6 Sol peaked at 20/23. Kimi K3 needed a two-hour extended budget just to reach the same score and cost $1.37 per solved task Grok 4.5 did not just win on accuracy. It won on economics too Grok’s coding efficiency is insane
Show more
Loading up the APIs for my next @VulcanBench run. Put a lot into this one, v3 eval suite with 23 new evals. Broader language coverage. And some really deep dives into making sure this doesn't have any of the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. Still, I'm new to benchmarking, so there's always a chance that I waste like $150 here. Hoping that's not the case, feels like I've tried to think of every edge case. This is the comparison I've been so excited to do: Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 All on Low and Medium effort, across real tasks that represent the work real engineering teams will be doing daily. No puzzles, no weird math problems, and looking not just at accuracy, but token use, cost, and time. In aggregate this will be a total of 138 runs, will wake up on Saturday to see the results.
Show more