đ¤¯ Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models?
... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents.
My July post ended with one BIG problem đ localmaxxers couldn't download the weights or load it into llama.cpp.â
Well... that part changed. đĨ
@TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL.
The new VL model adds:
đ Images
đĨ Video
đ¤ Visual/computer-use agents
đ§ 124B total / ~5.5B active
đ up to 1M context
âī¸ MIT
A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB
~1.7GB vision projector.
So roughly 81GB for a 124B multimodal MoE.
That puts it in range of these devices ...
đ§ 128GB Strix Halo
đ high-memory Apple Silicon
đŽ large/multi-GPU home rigs
â ī¸ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience.
But this is exactly the update I wanted back in July, weights on our own machines. đĨ
Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. đ
đ HF inclusionAI/Ling-3.0-flash-VL