登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Morgan
@morganlinton
Cofounder/CTO @BoldMetrics, Benchmarking the frontier @VulcanBench, Not an expert, always learning.
参加 January 2009
984 フォロー中    46.5K ファン
Thought I'd share a bit more about the models I'm using together to build my first cybersecurity eval suite for @VulcanBench Starting with Opus 4.8 Medium for planning and the suite mvp build, then sending to GPT 5.6 Sol Medium for updates/polishing, then sending to Grok 4.6 High to perfect and really make sure they're good. Three step process, and yes, as you probably noticed, Opus 4.8, not 5, and I only use High with one model. Still enjoy Opus 4.8 for prototyping and conversing with, it's a fun model to ideate with on eval suites. But, it's far from perfect. GPT 5.6 Sol adds quite a bit of polish, and catches stuff that me and Opus miss. In the end though, I trust Grok the most, and it tends to always think deeply about some aspect of the eval suite Opus and Sol both missed. And Grok is totally honest with me about how good or bad my current eval suite is, and pushes me to do better all the time. note: please don't confuse this with the stack I use for coding. this is a different stack than what I code with. I share my coding stack, and how it changes in my substack (
もっと見る