๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ivan Fioravanti แฏ…
@ivanfioravanti
GenAI/LLM addicted, Apple MLX, Cloud computing, Kubernetes, Technology Advisor, Investor and Co-Founder & Board Member of CoreView.
๊ฐ€์ž… June 2009
1.5K ํŒ”๋กœ์ž‰ ์ค‘    40.7K ํŒฌ
Hermes Agent as Master in a GDR (Blades in the Dark here) running locally on M3 Ultra using DeepSeek V4 Flash q4-imatrix with ds4 by @antirez Testing side by side with online version and apart from the speed, quality is nearly identical so far. I'll keep testing to see how far I can push with context! 1M goal ๐Ÿš€ Video below to show the decoding speed is still good at ~120K context. Pre fill is slow, but ok. I started with --kv-disk-space-mb 32756 to try with very large context. Image generation currently is with Grok Imagine, but I'm planning to update the skill to use mflux and a local model ๐Ÿ’ช Thanks @teomurgi for sharing this amazing idea! This is benchmarking while having fun!
๋” ๋ณด๊ธฐ