註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Joe Barrow
@barrowjoseph
NLP + Machine Learning Prev: NLP Researcher @AdobeResearch, PhD @ClipUmd.
加入 December 2012
517 正在關注    5K 粉絲
Good engineering is about observability. Speculative decoding accelerates LLM inference, but you're running it blind. Every rejected draft token is wasted compute, but are you looking at the drafts? This weekend I wrote specspecs to solve that!
顯示更多