ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅŠ ๅ…ฅ July 2023
549 ๆญฃๅœจๅ…ณๆณจ    11.2K ็ฒ‰ไธ
๐Ÿคฏ Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models? ... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents. My July post ended with one BIG problem ๐Ÿ‘‰ localmaxxers couldn't download the weights or load it into llama.cpp.โ€ Well... that part changed. ๐Ÿ”ฅ @TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL. The new VL model adds: ๐Ÿ‘€ Images ๐ŸŽฅ Video ๐Ÿค– Visual/computer-use agents ๐Ÿง  124B total / ~5.5B active ๐Ÿ“š up to 1M context โš–๏ธ MIT A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB ~1.7GB vision projector. So roughly 81GB for a 124B multimodal MoE. That puts it in range of these devices ... ๐Ÿง  128GB Strix Halo ๐ŸŽ high-memory Apple Silicon ๐ŸŽฎ large/multi-GPU home rigs โš ๏ธ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience. But this is exactly the update I wanted back in July, weights on our own machines. ๐Ÿ”ฅ Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. ๐Ÿ‘€ ๐Ÿ”— HF inclusionAI/Ling-3.0-flash-VL
ๆ˜พ็คบๆ›ดๅคš