가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Silicon Atlas
@Silicon_Atlas
Evidence-first AI semiconductor analysis: what new silicon claims prove, where bottlenecks move, and whether gains survive at system and economic scale.
가입 March 2021
123 팔로잉 중    2.9K 팬
A single AI answer combines two kinds of work: running the model and managing the service around it. In a typical GPU-based system, the CPU prepares inputs, schedules work, handles tool requests, and streams results. These jobs involve varied control flow and input/output. The GPU performs the model’s large, parallel calculations, applying similar operations across many pieces of data. That model work has two familiar phases. Prefill processes the prompt and builds the state needed for generation. Decode repeatedly uses the model to predict the next token, a piece of text. The weather example adds another handoff. The model generates a tool call; the application sends it to a weather service. Once the result returns, the model uses it to produce the answer. While that request waits, other requests can use the GPU. Now follow the orange line beneath the diagram. The KV cache stores attention state from previously processed tokens. If the server retains it during the tool call, it continues occupying memory while this request’s model computation pauses. Reusing that state can avoid recomputing the cached prefix when generation resumes. A faster GPU can shorten model computation. It cannot eliminate the wait for an external service. Scheduling and cache management therefore matter alongside raw compute speed. Two useful checks: Can your server schedule other requests during tool waits? And does it keep the KV cache, offload it, or discard it and recompute later?
더 보기