๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
This is a crazy, weird Qwen3.8-27B โ€œuncensoredโ€ release. It isn't because it's uncensored, but because the author measured what the guardrail surgery cost to get it uncensored and then preserved MTP. Jonathan Coletti took Qwen3.8-27B and used Heretic model surgery to reduce its refusal behavior. Measured on ~100 held-out test prompts ... ๐Ÿ›‘ Base Qwen3.8-27B โ†’ w/98 refusals ๐Ÿ”“ Modified model โ†’ w/only 12 refusals Meanwhile, the model's capability barely changed ... ๐Ÿง  MMLU 83.4 โ†’ 83.3 ๐ŸŽฏ HellaSwag 82.8 โ†’ 82.9 ๐Ÿ“Š Mean across 4 tests โ†’ ~0.5 point drop Abliteration initially dropped Qwen's MTP head and the author then grafted all 15 MTP tensors back from the original Qwen3.8-27B and then verified that the MTP block actually survived GGUF quantization. So the modified model you still get ... ๐Ÿง  27B dense Qwen ๐Ÿ‘๏ธ multimodal/vision ๐Ÿ“š 262K context architecture ๐Ÿ”ฎ native MTP speculative decoding ๐Ÿฆ™ llama.cpp GGUF And mradermacher already has iMatrix quantizations ๐Ÿ“ฆ IQ3_S โ†’ 12.7GB ๐Ÿ“ฆ IQ3_M โ†’ 12.9GB ๐Ÿ“ฆ Q3_K_M โ†’ 13.6GB ๐Ÿ‘€ ๐Ÿ“ฆ IQ4_XS โ†’ 15.4GB ๐Ÿ“ฆ Q4_K_M โ†’ 16.9GB That makes ~13GB versions especially interesting for 16GB GPU testing. ๐Ÿ”ฅ โš ๏ธ File size โ‰  VRAM requirement, so don't assume the 15.4GB IQ4 automatically fits comfortably in 16GB with useful context. Also, I have found that โ€œuncensoredโ€ model variants don't mean capability is proven unchanged. For example, their published comparison doesn't yet cover coding, math, multilingual, or vision quality. ๐Ÿ˜
๋” ๋ณด๊ธฐ