登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
参加 January 2009
756 フォロー中    43K ファン
Quick question @ArtificialAnlys, in this chart of cost per task, what effort level are you using with Fable 5? I've found that Fable Five Low and Medium effort can match the accuracy of most other frontier models on the same task for real engineering work. So I'm wondering if this chart gives an accurate cost per task if Fable 5 could have still performed well across these tasks at a lower effort level, but a higher effort level is used across tasks when it's not needed. Genuine curiousity!
もっと見る