登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Nathan
@nathanhabib1011
Evals @ huggingface 🤗
参加 December 2015
646 フォロー中    1.7K ファン
NEW BENCHMARK ON THE @huggingface HUB Your favorite LLM agent can write a script. Can it survive 300+ steps in a terminal without losing the plot? That's what Long-Horizon Terminal-Bench (LHTB) measures, 46 tasks, contamination-resistant, hidden verifiers that check real state, not vibes. 🏆 Leaderboard 🥇 @MiniMax_AI 's Minimax M3 🥈 @Kimi_Moonshot 's Kimi k2.7 Code 🥉 @Zai_org's GLM 5.2 Congrats to @zli12321 and team. Super excited to see how smaller local models perf on this benchmark @0xSero :)
もっと見る