注册并分享邀请链接,可获得视频播放与邀请奖励。

METR
@METR_Evals
We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
加入 September 2023
41 正在关注    54.9K 粉丝
How well can LLM agents complete diverse tasks compared to skilled humans? Our preliminary results indicate that our baseline agents based on several public models (Claude 3.5 Sonnet and GPT-4o) complete a proportion of tasks similar to what humans can do in ~30 minutes. 🧵
显示更多
0
10
426
95
转发到社区