가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Morgan
@morganlinton
Cofounder @BoldMetrics: the AI body data engine. Mad Scientist @VulcanBench: benchmarking models across effort levels on real coding tasks. Not an expert.
가입 January 2009
756 팔로잉 중    43K
Quick question @ArtificialAnlys, in this chart of cost per task, what effort level are you using with Fable 5? I've found that Fable Five Low and Medium effort can match the accuracy of most other frontier models on the same task for real engineering work. So I'm wondering if this chart gives an accurate cost per task if Fable 5 could have still performed well across these tasks at a lower effort level, but a higher effort level is used across tasks when it's not needed. Genuine curiousity!
더 보기