NEW BENCHMARK ON THE
@huggingface HUB
Your favorite LLM agent can write a script. Can it survive 300+ steps in a terminal without losing the plot?
That's what Long-Horizon Terminal-Bench (LHTB) measures, 46 tasks, contamination-resistant, hidden verifiers that check real state, not vibes.
๐ Leaderboard
๐ฅ
@MiniMax_AI 's Minimax M3
๐ฅ
@Kimi_Moonshot 's Kimi k2.7 Code
๐ฅ
@Zai_org's GLM 5.2
Congrats to
@zli12321 and team. Super excited to see how smaller local models perf on this benchmark
@0xSero :)