If you don’t know where to start, I’d highly recommend tensor economics, from @tugot17 and @felix_red_panda
First post here, but they’re all really good and worth working through:
Good engineering is about observability.
Speculative decoding accelerates LLM inference, but you're running it blind. Every rejected draft token is wasted compute, but are you looking at the drafts?
This weekend I wrote specspecs to solve that!