Run Ling 3.0 Flash locally ๐
We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4!
AD-Q5_K_M is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp default quant of the same size.