註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Nathan
@nathanhabib1011
Evals @ huggingface 🤗
加入 December 2015
646 正在關注    1.7K 粉絲
🦞 Claw-Eval 🦞 🥇 @XiaomiMiMo's MiMo-V2.5-Pro at 1T 🥈 @Zai_org GLM5.1 at 754B 🥉 @XiaomiMiMo MiMo-V2.5 at 310B Congrats to @XiaomiMiMo for having 2 models in the top 3! The most impressive result though is @deepseek_ai with DeepSeek v4 flash a 210B model on par with models 4 times its size.. Super interesting bench from @_TobiasLee and team. Gathering tasks from @openclaw, PinchBench, OfficeQA, OneMillion-Bench, Finance Agent, and Terminal-Bench 2.0. This might be one of the more interesting and useful benchmark right now. Whether you are using @NousResearch's Hermes agent or @steipete's OpenClaw, you should choose your model according to real world tasks.
顯示更多
0
6
145
16
轉發到社區