注册并分享邀请链接,可获得视频播放与邀请奖励。

Morgan
@morganlinton
Cofounder/CTO @BoldMetrics, Benchmarking the frontier @VulcanBench, Not an expert, always learning.
加入 January 2009
984 正在关注    46.5K 粉丝
Thought I'd share a bit more about the models I'm using together to build my first cybersecurity eval suite for @VulcanBench Starting with Opus 4.8 Medium for planning and the suite mvp build, then sending to GPT 5.6 Sol Medium for updates/polishing, then sending to Grok 4.6 High to perfect and really make sure they're good. Three step process, and yes, as you probably noticed, Opus 4.8, not 5, and I only use High with one model. Still enjoy Opus 4.8 for prototyping and conversing with, it's a fun model to ideate with on eval suites. But, it's far from perfect. GPT 5.6 Sol adds quite a bit of polish, and catches stuff that me and Opus miss. In the end though, I trust Grok the most, and it tends to always think deeply about some aspect of the eval suite Opus and Sol both missed. And Grok is totally honest with me about how good or bad my current eval suite is, and pushes me to do better all the time. note: please don't confuse this with the stack I use for coding. this is a different stack than what I code with. I share my coding stack, and how it changes in my substack (
显示更多