🤯 Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models?
... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents.
My July post ended with one BIG problem 👉 localmaxxers couldn't download the weights or load it into llama.cpp.”
Well... that part changed. 🔥
@TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL.
The new VL model adds:
👀 Images
🎥 Video
🤖 Visual/computer-use agents
🧠 124B total / ~5.5B active
📚 up to 1M context
⚖️ MIT
A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB
~1.7GB vision projector.
So roughly 81GB for a 124B multimodal MoE.
That puts it in range of these devices ...
🧠 128GB Strix Halo
🍎 high-memory Apple Silicon
🎮 large/multi-GPU home rigs
⚠️ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience.
But this is exactly the update I wanted back in July, weights on our own machines. 🔥
Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. 👀
🔗 HF inclusionAI/Ling-3.0-flash-VL