how inference engines actually work - releasing full talk slides!
i cover everything in the lifetime of a request e2e:
> the inference engine
> kv + prefix caching
> continuous batching
> paged attention
> chunked prefill
> sampling
> agentic loops from inside the engine