็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅ‚ๅŠ  July 2023
549 ใƒ•ใ‚ฉใƒญใƒผไธญ    11.2K ใƒ•ใ‚กใƒณ
Okay, so specialized decision models are becoming the rage. Bespoke Labs just released Nimble, an open 9B decision model built from Qwen3.5-9B. And apparently they built it in ONE DAY. The recipe: ๐Ÿง  Qwen3.5-9B ๐ŸŽฏ LoRA fine-tune ๐Ÿ“š only 2,676 training examples โŒ no Jev distillation โŒ no RL โœ… open model โœ… open data โœ… open training recipe Instead of generating explanations or JSON, Nimble scores the allowed answers: YES / NO โ†’ A / B / C โ†’ route 1 / 2 / 3 โ†’ severity 1โ€“5 Then returns the choice + probabilities. Example of Bespoke's 324-example held-out test: Qwen3.5-9B โ†’ 66.36% Qwen3.8-27B โ†’ 84.88% ๐Ÿ”ฅ Nimble-9B โ†’ 90.12% Jev โ†’ 93.21% And because it isn't generating a bunch of tokens: โšก ~106 ms median on H100 ๐ŸŽ ~444 ms median on an M5 Pro 64GB ๐Ÿ‘ˆ ๐Ÿ‘€ ๐ŸŽฏ So this is the key, giant reasoning model doesn't need to answer every tiny agent decision. Let the big model think, then let a small local model handle the ... ๐Ÿ› ๏ธ tool selection ๐Ÿงญ routing โœ… verification ๐Ÿšจ policy decisions ๐Ÿ“Š scoring ๐Ÿ”„ retry / stop Fast. Local. No API call. โš ๏ธ Bespoke says this is a narrow synthetic 324-example evaluation, NOT a standardized System One benchmark. ๐Ÿ”— GH: /bespokelabsai/nimble
ใ‚‚ใฃใจ่ฆ‹ใ‚‹