Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Seeing a high volume of posts on X about Astra being nerfed.
Putting together a plan to test this with VulcanBench to move from anecdotal evidence to a more data-driven approach.
Since I benchmarked it launch week, I do have data to compare how it performed at launch vs. now.
But VulcanBench only looks at coding, not things like 3d rendering, game design, etc. and a lot of the nerf reports seemed to be focused on this so I might be somewhat limited.
Still, worth seeing if we can notice any kind of difference.
Let's let the data do the talking and answer the question: is Astra really nerfed?