가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
가입 May 2018
2.1K 팔로잉 중    12.6K 팬
Is your CPU:GPU ratio still built for training? In agentic workloads, 50-90% of end-to-end latency is CPU-side tool processing, not GPU math (Georgia Tech + Intel). That's pushing the CPU:GPU ratio from 1:8 in training toward 1:1, sometimes 4:1. And vLLM's CPU backend already runs PagedAttention, prefix caching, and continuous batching across x86, Arm, IBM Z, and experimental Apple Silicon. @_soyr_ and @__gracecaroline on where inference compute should live:
더 보기