Register and share your invite link to earn from video plays and referrals.

Nathan
@nathanhabib1011
Evals @ huggingface ๐Ÿค—
646 Following    1.7K Followers
NEW BENCHMARK ON THE @huggingface HUB Your favorite LLM agent can write a script. Can it survive 300+ steps in a terminal without losing the plot? That's what Long-Horizon Terminal-Bench (LHTB) measures, 46 tasks, contamination-resistant, hidden verifiers that check real state, not vibes. ๐Ÿ† Leaderboard ๐Ÿฅ‡ @MiniMax_AI 's Minimax M3 ๐Ÿฅˆ @Kimi_Moonshot 's Kimi k2.7 Code ๐Ÿฅ‰ @Zai_org's GLM 5.2 Congrats to @zli12321 and team. Super excited to see how smaller local models perf on this benchmark @0xSero :)
Show more
๐Ÿšจ ITS HERE๐Ÿšจ โšกQwen3.8-27B - Multimodal dense model. With just 27B parameters, it outperforms most models on the @huggingface leaderboards.. - Built for builders: Apache 2.0 license Swe-Bench Pro, Deep-Swe and Claw-Eval, those 3 benchmarks evaluate how good a model is at being an agent, needless to say @Alibaba_Qwen team delivered..
Show more
while we cannot use fable 5 and gpt 5.6 will be screened before allowed to use the model, open source model builders are hard at work to close the gap. Ornith-1.0-397B from @ornith_, the best coding open source model.
Show more
๐Ÿฆž Claw-Eval ๐Ÿฆž ๐Ÿฅ‡ @XiaomiMiMo's MiMo-V2.5-Pro at 1T ๐Ÿฅˆ @Zai_org GLM5.1 at 754B ๐Ÿฅ‰ @XiaomiMiMo MiMo-V2.5 at 310B Congrats to @XiaomiMiMo for having 2 models in the top 3! The most impressive result though is @deepseek_ai with DeepSeek v4 flash a 210B model on par with models 4 times its size.. Super interesting bench from @_TobiasLee and team. Gathering tasks from @openclaw, PinchBench, OfficeQA, OneMillion-Bench, Finance Agent, and Terminal-Bench 2.0. This might be one of the more interesting and useful benchmark right now. Whether you are using @NousResearch's Hermes agent or @steipete's OpenClaw, you should choose your model according to real world tasks.
Show more