註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Ivan Fioravanti ᯅ
@ivanfioravanti
GenAI/LLM addicted, Apple MLX, Cloud computing, Kubernetes, Technology Advisor, Investor and Co-Founder & Board Member of CoreView.
加入 June 2009
1.5K 正在關注    40.7K 粉絲
Hermes Agent as Master in a GDR (Blades in the Dark here) running locally on M3 Ultra using DeepSeek V4 Flash q4-imatrix with ds4 by @antirez Testing side by side with online version and apart from the speed, quality is nearly identical so far. I'll keep testing to see how far I can push with context! 1M goal 🚀 Video below to show the decoding speed is still good at ~120K context. Pre fill is slow, but ok. I started with --kv-disk-space-mb 32756 to try with very large context. Image generation currently is with Grok Imagine, but I'm planning to update the skill to use mflux and a local model 💪 Thanks @teomurgi for sharing this amazing idea! This is benchmarking while having fun!
顯示更多
0
11
71
5
轉發到社區