็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅ‚ๅŠ  July 2023
549 ใƒ•ใ‚ฉใƒญใƒผไธญ    11.2K ใƒ•ใ‚กใƒณ
๐Ÿ’ฅ Okay, yesterday we had MiniCPM5-2B, but today @OpenSquilla turned Qwen3.5-9B into a ~5.7GB LOCAL agent model, and their training method is interesting. The new NeoHorse-1-9B family is a post-trained Qwen3.5-9B specifically for agents, tool use + coding. It learned from actual multi-model agent execution traces. Setup ... ๐Ÿ”ง tool calls โŒ failed attempts ๐Ÿ”€ model handoffs ๐Ÿง  changed plans โœ… environment-verified successes Then they fed those experiences back into the 9B model. Their reported results vs Qwen3.5-9B: ๐Ÿค– QwenClaw: 44.04 โ†’ 48.73 ๐Ÿ“Œ PinchBench: 74.55 โ†’ 82.25 ๐Ÿ”„ VitaBench: 31.25 โ†’ 42.25 ๐Ÿ›  BFCL v4: 64.88 โ†’ 67.43 ๐Ÿ’ป HumanEval: 92.68 โ†’ 98.17 Overall: 65.60 โ†’ 69.04 This is the model I'm most interested in for ThumbLLM ๐Ÿง  ~9B parameters ๐Ÿ’พ Q4_K_M GGUF ~5.7GB ๐Ÿฆ™ llama.cpp ๐ŸŸข Ollama ๐Ÿ–ฅ๏ธ LM Studio ๐Ÿค– Hermes / OpenClaw ๐Ÿ“š 262K native context โš–๏ธ Apache 2.0 So you can run this entirely locally on CPU or GPU. This isn't a bigger model. It's an attempt to make the same small local model better at actually DOING things. ๐Ÿ”ฅ โš ๏ธ Benchmarks are reported by the NeoHorse team and I want to do these tests myself next.
ใ‚‚ใฃใจ่ฆ‹ใ‚‹