Today’s video is on DFlash 2 from Inco ai and speculative decoding!
I explain what DFlash is, how it speeds up local models, and then run an experiment with three versions of a local model, one with no speculative decoding, one with DFlash, and one with DFlash 2 to see the difference! Check it out!
@zhijianliu_
Tinkering with DFlash2: How to Speed Up Local AI Models