Fable quality. We will soon release a new quantized inference framework that will let you run this on a single GPU (B300) with ~50 tok/s in good quality. Local Fable LFG!
DeepSeek silently released V4-Pro 0813, up 15.8% on Terminal Bench from their April Preview model, with Fable 5 performance at ~57x cheaper cost.
1.6T param, 49B active, 1M context. This is the best price-to-perfomance model on the market right now.
Available in ClinePass now!