Register and share your invite link to earn from video plays and referrals.

Jesse Zhang
@thejessezhang
Joined October 2021
203 Following    102.1K Followers
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
Show more
We were very excited by @inco_ai's Dflash2 release and started building with it immediately. A few tweaks to the training recipe further improved acceptance length by another 33%. Full write-up below.
Show more