๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Mia
@MiaAI_lab
Building with AI & LLMs | Insights, recipes, tools & honest experiments.
๊ฐ€์ž… July 2022
413 ํŒ”๋กœ์ž‰ ์ค‘    36.1K ํŒฌ
TensorFold is going to change everything ๐Ÿš€ Exciting times ahead! โณ
TensorFold Inference Engine is here ๐Ÿš€ I spent six months making one weight read count for more than one token on Apple Silicon. Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes. Qwen 3.8 27B MLX 4Bit - 120-124tks Nemotron Lightning MLX 4Bit - 188-206tks Qwen3.8 Flash Next MLX 4Bit - 88-92tks CUDA Implementation is in Alpha showing strong gains. The Repo is in the comments ๐Ÿ‘‡๐Ÿผ
๋” ๋ณด๊ธฐ