TensorFold Inference Engine is here ๐
I spent six months making one weight read count for more than one token on Apple Silicon.
Draft tokens run through parallel lanes; the model verifies them together and keeps only what passes.
Qwen 3.8 27B MLX 4Bit - 120-124tks
Nemotron Lightning MLX 4Bit - 188-206tks
Qwen3.8 Flash Next MLX 4Bit - 88-92tks
CUDA Implementation is in Alpha showing strong gains.
The Repo is in the comments ๐๐ผ