註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

METR
@METR_Evals
We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.
加入 September 2023
41 正在關注    54.9K 粉絲
How well can LLM agents complete diverse tasks compared to skilled humans? Our preliminary results indicate that our baseline agents based on several public models (Claude 3.5 Sonnet and GPT-4o) complete a proportion of tasks similar to what humans can do in ~30 minutes. 🧵
顯示更多
0
10
426
95
轉發到社區