๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
Okay, so specialized decision models are becoming the rage. Bespoke Labs just released Nimble, an open 9B decision model built from Qwen3.5-9B. And apparently they built it in ONE DAY. The recipe: ๐Ÿง  Qwen3.5-9B ๐ŸŽฏ LoRA fine-tune ๐Ÿ“š only 2,676 training examples โŒ no Jev distillation โŒ no RL โœ… open model โœ… open data โœ… open training recipe Instead of generating explanations or JSON, Nimble scores the allowed answers: YES / NO โ†’ A / B / C โ†’ route 1 / 2 / 3 โ†’ severity 1โ€“5 Then returns the choice + probabilities. Example of Bespoke's 324-example held-out test: Qwen3.5-9B โ†’ 66.36% Qwen3.8-27B โ†’ 84.88% ๐Ÿ”ฅ Nimble-9B โ†’ 90.12% Jev โ†’ 93.21% And because it isn't generating a bunch of tokens: โšก ~106 ms median on H100 ๐ŸŽ ~444 ms median on an M5 Pro 64GB ๐Ÿ‘ˆ ๐Ÿ‘€ ๐ŸŽฏ So this is the key, giant reasoning model doesn't need to answer every tiny agent decision. Let the big model think, then let a small local model handle the ... ๐Ÿ› ๏ธ tool selection ๐Ÿงญ routing โœ… verification ๐Ÿšจ policy decisions ๐Ÿ“Š scoring ๐Ÿ”„ retry / stop Fast. Local. No API call. โš ๏ธ Bespoke says this is a narrow synthetic 324-example evaluation, NOT a standardized System One benchmark. ๐Ÿ”— GH: /bespokelabsai/nimble
๋” ๋ณด๊ธฐ