가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Jesse Zhang
@thejessezhang
가입 October 2021
203 팔로잉 중    102.1K 팬
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
더 보기
We were very excited by @inco_ai's Dflash2 release and started building with it immediately. A few tweaks to the training recipe further improved acceptance length by another 33%. Full write-up below.
더 보기