Hermes Agent as Master in a GDR (Blades in the Dark here) running locally on M3 Ultra using DeepSeek V4 Flash q4-imatrix with ds4 by
@antirez
Testing side by side with online version and apart from the speed, quality is nearly identical so far.
I'll keep testing to see how far I can push with context!
1M goal 🚀 Video below to show the decoding speed is still good at ~120K context. Pre fill is slow, but ok. I started with --kv-disk-space-mb 32756 to try with very large context.
Image generation currently is with Grok Imagine, but I'm planning to update the skill to use mflux and a local model 💪
Thanks
@teomurgi for sharing this amazing idea! This is benchmarking while having fun!