Register and share your invite link to earn from video plays and referrals.

Search results for VUCCA
VUCCA community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including VUCCA
Grok 4.5 just ranked #1# on VulcanBench, outperforming Claude Fable 5, GPT-5.6 Sol and Kimi K3 across 23 real-world merged PRs in Python, Rust, TypeScript, JavaScript and Go Insane efficiency: • 21/23 tasks solved - 91.3% • Best result across all 11 configurations tested • Just $0.32 per solved task Claude Fable 5 and GPT-5.6 Sol peaked at 20/23. Kimi K3 needed a two-hour extended budget just to reach the same score and cost $1.37 per solved task Grok 4.5 did not just win on accuracy. It won on economics too Grok’s coding efficiency is insane
Show more
Loading up the APIs for my next @VulcanBench run. Put a lot into this one, v3 eval suite with 23 new evals. Broader language coverage. And some really deep dives into making sure this doesn't have any of the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. Still, I'm new to benchmarking, so there's always a chance that I waste like $150 here. Hoping that's not the case, feels like I've tried to think of every edge case. This is the comparison I've been so excited to do: Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5 All on Low and Medium effort, across real tasks that represent the work real engineering teams will be doing daily. No puzzles, no weird math problems, and looking not just at accuracy, but token use, cost, and time. In aggregate this will be a total of 138 runs, will wake up on Saturday to see the results.
Show more
Okay, just saw the @VulcanBench results for Grok 4.6 across all effort levels, and this might be the most interesting benchmark I have ever done. This is one pass, think I'll probably need to do three passes, but I'll share the results in the morning. Kinda need to see if someone at @SpaceXAI might be able to hook me up with some credits though. It'll cost me around $150 for me to do another two passes, which I'll totally do, but feels like I need to see if there's a time to ask, it would be now. So...if anyone knows someone that might have some power to help an independent benchmarking nerd like me, please send them my way 🖖
Show more
BREAKING: Grok Voice Think Fast 2.0 outperforms GPT Realtime in VulcanBench's Voice Tax benchmark. • Achieved 99.0% accuracy on text. • Scored 95.7% on audio across the same 200-question benchmark. • Completed voice interactions nearly 2× faster than GPT Realtime.
Show more
Okay, my new 23-task eval suite for @VulcanBench is done. A lot of little details I wanted to get right with this. It was important to me that I had more broad language coverage, and also that it takes into account the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes. I'm new to benchmarking, but learning more every time I create a new set of evals. I'm calling this set v3, getting ready to run the first smoke tests. Then this weekend, phew, safe to say I have a lot of models to test! I will be committing these to Github so anyone can take a look and give me feedback on these as well. Live long and Benchmark 🖖
Show more
Well it finally happened. A lab gave me credits to use for benchmarking with @VulcanBench so I don't have to move out of my house and into a paper box in order to afford these benchmarks 😅 📦 Huge thanks to @cursor_ai, who has now made it possible for me to run a 3 pass benchmark with Grok 4.6, and likely whatever else comes out this week, without breaking the bank 🙏 Proud to bring independent benchmarking to the world, with 100% real engineering tasks, the stuff actual engineering teams would do with a model. No puzzles, no exotic architecture. And all open source, so every eval, and the benchmarking software itself is all available for you to review, critique, fork, etc. Live long and benchmark 🖖
Show more
I have now had more than a few companies reach out to me about doing stack-specific benchmarks and model reports with @VulcanBench. This wasn't something I thought of, but then it kinda clicked. I'm also hearing now more and more about challenges eng leaders are having taking all the benchmarks that are out there, and figuring out how to apply them to their teams, and what kind of model routing might be best for them based on their stack and codebase size. This is also where effort levels really matter. I've heard from a lot of people that, "my team just uses one model, high effort, for everything." And yeah, that's not optimal. It's both expensive, and slow, and you'll end up with your eng team waiting around for hours, when they could get a task done faster, and cheaper, if they had better model routing. This is different from all the model routers companies are coming out with, because those are general, with VulcanBench, I can get really specific using actual PRs from real work that eng teams are doing to customize model routing just for them. And, since I now have a growing list of companies that want this, I guess it just makes sense to add it to the VulcanBench site and start offering it as a service, and we'll just see where it goes! No idea how to price it, just putting it there as an email link and will take it convo by convo. You can read more about it here:
Show more
Rihanna begs ‘RHOBH’ star not to quit series as Season 16 cast rumors swirl: ‘We need more of you’