登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Silicon Atlas
@Silicon_Atlas
Evidence-first AI semiconductor analysis: what new silicon claims prove, where bottlenecks move, and whether gains survive at system and economic scale.
参加 March 2021
123 フォロー中    2.9K ファン
Why does waiting cut the price in half? If you need an answer right now, you pay the standard rate. If the work can wait, providers such as OpenAI and Anthropic will run it through their batch APIs at half price. Same model. Same task. A more flexible deadline. So why should timing change the price that much? Waiting gives the provider room to schedule. Your job can be placed wherever there is spare capacity, instead of competing for it at the moment you asked. Summarizing yesterday's support tickets can wait. A reply in a live chat cannot. There is also a hardware effect underneath. To produce each token, the model reads the data that defines it, its weights, out of memory. That read is expensive, and it happens whether one request or many are being served. When requests run together, they share it, and the hardware gets more useful work out of the same data movement. That is batching: running several requests together so they share some of the overhead. The saving has limits. Each request still needs its own computation and its own memory for its own conversation, and a group eventually runs into both. The 50 percent is a pricing decision, not a direct measure of the hardware saving. But the direction is real. Your willingness to wait gives the provider more ways to make the work cheap.
もっと見る