Is your CPU:GPU ratio still built for training? In agentic workloads, 50-90% of end-to-end latency is CPU-side tool processing, not GPU math (Georgia Tech + Intel).
That's pushing the CPU:GPU ratio from 1:8 in training toward 1:1, sometimes 4:1. And vLLM's CPU backend already runs PagedAttention, prefix caching, and continuous batching across x86, Arm, IBM Z, and experimental Apple Silicon.
@_soyr_ and
@__gracecaroline on where inference compute should live: