註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Jesse Zhang
@thejessezhang
加入 October 2021
203 正在關注    102.1K 粉絲
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
顯示更多
We were very excited by @inco_ai's Dflash2 release and started building with it immediately. A few tweaks to the training recipe further improved acceptance length by another 33%. Full write-up below.
顯示更多