My First Model On
@huggingface Qwen 3.8 27B Unleashed HIT 105,666 Downloads ๐ฅ
Qwen3.8-27B Unleashed UD-Q3_K_XL hits near 100 tk/s @ 250k ctx on 1 x 4090.
- DFlash2
- LoopSpec
- our async-Q4 Ada kernel
- Q4 KV cache
- full GPU offload
Context Decode Speed
โโโโโโโโโ
Short 134 tok/s
โโโโโโโโโ
8K 133 tok/s
โโโโโโโโโ
64K 131 tok/s
โโโโโโโโโ
128K 109 tok/s
โโโโโโโโโ
250K 92 tok/s
In the actual Hermes harness:
- ~24K context: 98.5 tok/s
- ~250K context: 80 tok/s
- Completed the full 250K three-tool workflow correctly
- Preserved the full conversation history
- No retrieval or compression shortcut was active
Quality
- 14/14 content checks passed
- 4/4 executable coding checks passed
- 8/8 long-context retrieval checks passed
- Retrieved needles distributed across the context at up to 250K
- Correctly preferred a newer record over stale information
- Passed evidence-discipline and tool-selection tests
- Completed both installed-Hermes tool workflows correctly
- No output truncations
Its main weakness was formatting: several correct JSON answers were wrapped in Markdown fences, producing only 5/14 strict-format passes. Thatโs easy to address at the harness or prompt layer and wasnโt a reasoning failure.
Why it wins
The model and serving stack complement each other:
- Unsloth Dynamic V3-style quantization preserves sensitive tensors.
- Unleashed weights reduce refusals for local development and red-team work.
- DFlash2 drafts target-specific token blocks.
- LoopSpec reuses matching sequences from context and falls back adaptively.
- The async-Q4 kernel improves Ada execution.
- Q4 KV cache makes 262K context practical on 24GB VRAM.
Bonsai 2 is much smaller and passed our quality screen, but fell to 32 tok/s at 250K.
Emperoโs MoE processed cold prompts quickly but decoded around 53 tok/s in the direct 250K test.
Our winner sustained 92 tok/s in the matched direct test.