One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass.
It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client.
This is called "speculative decoding" and lets you generate faster while maintaining the output quality.
Great write-up from the
@DecagonAI team on how we've been doing this: