Thought I'd share a bit more about the models I'm using together to build my first cybersecurity eval suite for
@VulcanBench
Starting with Opus 4.8 Medium for planning and the suite mvp build, then sending to GPT 5.6 Sol Medium for updates/polishing, then sending to Grok 4.6 High to perfect and really make sure they're good.
Three step process, and yes, as you probably noticed, Opus 4.8, not 5, and I only use High with one model. Still enjoy Opus 4.8 for prototyping and conversing with, it's a fun model to ideate with on eval suites.
But, it's far from perfect. GPT 5.6 Sol adds quite a bit of polish, and catches stuff that me and Opus miss.
In the end though, I trust Grok the most, and it tends to always think deeply about some aspect of the eval suite Opus and Sol both missed. And Grok is totally honest with me about how good or bad my current eval suite is, and pushes me to do better all the time.
note: please don't confuse this with the stack I use for coding. this is a different stack than what I code with. I share my coding stack, and how it changes in my substack (