๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Nathan
@nathanhabib1011
Evals @ huggingface ๐Ÿค—
๊ฐ€์ž… December 2015
646 ํŒ”๋กœ์ž‰ ์ค‘    1.7K ํŒฌ
NEW BENCHMARK ON THE @huggingface HUB Your favorite LLM agent can write a script. Can it survive 300+ steps in a terminal without losing the plot? That's what Long-Horizon Terminal-Bench (LHTB) measures, 46 tasks, contamination-resistant, hidden verifiers that check real state, not vibes. ๐Ÿ† Leaderboard ๐Ÿฅ‡ @MiniMax_AI 's Minimax M3 ๐Ÿฅˆ @Kimi_Moonshot 's Kimi k2.7 Code ๐Ÿฅ‰ @Zai_org's GLM 5.2 Congrats to @zli12321 and team. Super excited to see how smaller local models perf on this benchmark @0xSero :)
๋” ๋ณด๊ธฐ