I kept tweaking the setup and running tests, and the results turned out even better than I expected. It feels ready for a full rollout at work now.
That's a big relief when working with sensitive data. I don't have to worry about it being sold or exposed in a hack.
The only catch is speed. It's a little slow on my AMD GPU.
Got Qwen3.8 running locally and made a particle visualization for it.
The model runs in LM Studio on my PC with a 20GB RX 7900 XT. I work from my Mac, and Pi Agent talks to it over the local network.
Current setup:
- unsloth/Qwen3.8-27B-UD-Q3_K_XL
- 128K context length
- 24 offloaded to GPU, automatically set by LM Studio
- Flash Attention on
- Q4_0 K/V cache on GPU
- Batch size 256, one concurrent prediction
- Default RoPE
- Speculative decoding and MTP off