๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
544 ํŒ”๋กœ์ž‰ ์ค‘    8.8K ํŒฌ
๐Ÿš€ Holy ๐Ÿ’ฉ! Major Local AI Breakthrough! ๐Ÿง  744B-parameter GLM-5.2 (1.5 TB total) is now running on just ~25 GB RAM โ€” no discrete GPU required! ๐Ÿ‘€ Wut!? Must test!! Italian engineer @JustVugg built Colibrรฌ, a pure C inference engine (single ~2.4k line file, zero runtime deps) that: โ€ข ๐Ÿ›ก๏ธ Keeps the dense core (~10 GB at int4) resident in RAM โ€ข ๐Ÿ“€ Streams 21,504+ MoE experts from fast NVMe on demand (only ~40B active per token) โ€ข โšก Supports native MTP speculative decoding + MLA attention Result: Frontier-class model on everyday consumer hardware! ๐Ÿ“Š Current speeds: โ€ข 25 GB RAM setup โ†’ 0.05โ€“0.1 tok/s (disk-bound) โ€ข Higher RAM + fast SSD โ†’ up to 1+ tok/s (warm) โ€ข It's a start... what could you do with a 5090? ๐Ÿ’ก Big opportunity: Pair it with Phison aiDAPTIV+ AI SSDs to kill the I/O bottleneck ๐Ÿ‘‰ smarter caching, prefetching & KV offload could make it dramatically faster! This is a huge step toward truly accessible local frontier AI. ๐Ÿ”— GitHub: JustVugg/colibri
๋” ๋ณด๊ธฐ