가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Karim Mattar
@MattarARK
AI Research Associate @ARKInvest | AI, Cloud & Semis | Research x Automation | | Disclosure:
가입 September 2024
250 팔로잉 중    6.8K 팬
I've been turning effort dials like they're free upgrades. They're not. Anthropic gets most of the juice early. OpenAI keeps climbing if you pay for max. So the default seat and the burst seat might not be the same model anymore. That changes how you build agents more than another leaderboard screenshot.
더 보기
There is something clearly different in how Anthropic & OpenAI scale effort levels. Doesn't show up on every benchmark, but these results on FrontierCode make it really clear. Anthropic models peak at lower effort levels, where as OpenAI models start low and climb up fairly consistently. The result is a better score at a lower cost for Anthropic models, but an unintuitive experience where increasing effort does not increase scores and might actually degrade performance (on this benchmark at least).
더 보기