yesterday we made them more compressed! today we make them faster than ever with speculative decoding!
up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe.
You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own
@tugot17
enjoy ๐ข๐