๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ash Hart
@ashxhart
Doing my bit to bring local AI to everyone | | | | Thoughts are my own.
๊ฐ€์ž… October 2014
375 ํŒ”๋กœ์ž‰ ์ค‘    5.2K ํŒฌ
TensorFold Inference Engine is here ๐Ÿš€ I spent six months making one weight read count for more than one token on Apple Silicon. Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes. Qwen 3.8 27B MLX 4Bit - 120-124tks Nemotron Lightning MLX 4Bit - 188-206tks Qwen3.8 Flash Next MLX 4Bit - 88-92tks CUDA Implementation is in Alpha showing strong gains. The Repo is in the comments ๐Ÿ‘‡๐Ÿผ
๋” ๋ณด๊ธฐ