註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

JJ Englert
@JJEnglert
I help non-technical people do more with AI. Enablement Lead at @Tenex_labs (Best AI Consulting Firm) 😎 Join our community for top 1% AI builders 👇🏼
加入 June 2011
3.8K 正在關注    36.9K 粉絲
I couldn't believe the results of this benchmarking test. I blind-tested 6 frontier models — Sonnet 5, Opus 4.8, Fable 5, GPT 5.5, GLM 5.2, and Kimi 2.5 — on knowledge work, not engineering. Writing. Analysis. Design. Vibe coding. Subagent simulations. The stuff we actually do all day. Same prompts, each model running its own subagent, every output graded blind before the labels were revealed. The results: → GLM 5.2 (the dirt-cheap open-source model from China) took first → Sonnet 5 took second, beating Fable 5 for knowledge work → Opus 4.8 got whooped Then I had Composer 2.5 judge every output independently — and it picked Fable 5 almost across the board. My blind picks and the AI judge completely disagreed. Full breakdown, method, and prompts in the video below. Which model are you running as your daily driver right now?
顯示更多