注册并分享邀请链接,可获得视频播放与邀请奖励。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
加入 May 2018
2.1K 正在关注    12.6K 粉丝
Is your CPU:GPU ratio still built for training? In agentic workloads, 50-90% of end-to-end latency is CPU-side tool processing, not GPU math (Georgia Tech + Intel). That's pushing the CPU:GPU ratio from 1:8 in training toward 1:1, sometimes 4:1. And vLLM's CPU backend already runs PagedAttention, prefix caching, and continuous batching across x86, Arm, IBM Z, and experimental Apple Silicon. @_soyr_ and @__gracecaroline on where inference compute should live:
显示更多