่จปๅ†Šไธฆๅˆ†ไบซ้‚€่ซ‹้€ฃ็ต๏ผŒๅฏ็ฒๅพ—ๅฝฑ็‰‡ๆ’ญๆ”พ่ˆ‡้‚€่ซ‹็Žๅ‹ตใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅŠ ๅ…ฅ July 2023
549 ๆญฃๅœจ้—œๆณจ    11.2K ็ฒ‰็ตฒ
๐Ÿคฏ Remember back in July I called Ling-3.0-flash one of the more interesting Local AI models? ... because Ling-3.0-flash combined 124B-scale capacity with only ~5B active parameters per token, strong frontierish benchmarks, and a design that looked promising for local agents. My July post ended with one BIG problem ๐Ÿ‘‰ localmaxxers couldn't download the weights or load it into llama.cpp.โ€ Well... that part changed. ๐Ÿ”ฅ @TheInclusionAI has now released the weights under MIT, an official GGUF exists for Ling-3.0-flash, and they've just released Ling-3.0-flash-VL. The new VL model adds: ๐Ÿ‘€ Images ๐ŸŽฅ Video ๐Ÿค– Visual/computer-use agents ๐Ÿง  124B total / ~5.5B active ๐Ÿ“š up to 1M context โš–๏ธ MIT A community Ling-3.0-flash-VL Q4_K_M GGUF is already ~79.3GB ~1.7GB vision projector. So roughly 81GB for a 124B multimodal MoE. That puts it in range of these devices ... ๐Ÿง  128GB Strix Halo ๐ŸŽ high-memory Apple Silicon ๐ŸŽฎ large/multi-GPU home rigs โš ๏ธ The VL GGUF is still experimental and currently uses a patched llama.cpp, so this is not yet a clean stock-llama.cpp experience. But this is exactly the update I wanted back in July, weights on our own machines. ๐Ÿ”ฅ Now somebody run that ~81GB VL stack on a 128GB Strix Halo and give us the tps. ๐Ÿ‘€ ๐Ÿ”— HF inclusionAI/Ling-3.0-flash-VL
้กฏ็คบๆ›ดๅคš