Run GGUF models directly with transformers. This work brings ggml's Metal kernels to the transformers ecosystem, increasing compatibility and performance. More info below
Millions of GGUF downloads later, those same llama.cpp checkpoints can now run in 🤗 transformers.
Same models, more ways to use them, and fast local inference on Mac powered by ggml kernels!
Blog:
ggml kernels: