Live now: our Inference Engineering Track from AI Engineer World's Fair 2026.
A benchmark tool told to run 200 queries a second that ran 38 and reported 200. A model that answered one prompt in a thousand with confident gibberish. A paper that dented memory chip stocks for a minute.
- Operating Distributed Inference Systems at Scale: Nishant Gupta & Naman Ahuja, Meta
- Routing LLM Inference in Production: From Engine Signals to Policy: Qianru Lao & Lu Zhang, OpenAI
- Are LLM Performance Benchmarks Reliable?: Ashok Chandrasekar & Jason Kramberger, Google
- Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads: Sitanshu Gupta, CoreWeave
- What's New in Inference Engineering:
@philip_kiely, Baseten
- Large clusters for small models:
@svonava, Superlinked
- The Frontier AI Inference Cloud for Agents: Byung-Gon (Gon) Chun, FriendliAI
- KV Cache-Aware Routing and P/D Disaggregation on Kubernetes: Yuchen Fama & Ashish Kamra, Red Hat
- Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story: Asaf Gardin &
@yuvalinthedeep, AI21
- Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards:
@f_makraduli, Superlinked