註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

shmidt
@shmidtqq
predicting the future with code / ai × finance / tg: shmidtqq
加入 October 2021
899 正在關注    10.9K 粉絲
There is one line in your system prompt that multiplies your bill by 5. datetime. now() The cache reads your prompt left to right and stops at the first difference. A date on line three means the 60,000 tokens of documentation below it get recomputed. Every request. Forever. The same million input tokens: $3.00 if the start of the prompt changed $0.30 if it matched byte for byte Not a discount. Not an enterprise deal. An official line on the pricing page: Input Price (Cache Hit). One agent run, 30 steps: miss → $6.08 hit → $1.22 1,000 runs a day → $1.77M a year apart. On one line of text. Moonshot themselves run 90%+ cache hit on coding traffic. That is their operating norm, not a blog aspiration. If you are at 30%, it is not the provider. 30 seconds, right now: grep -rn "now()\|uuid4()\|random" your_prompt_builder.py Everything it finds above your documents is billed at 10x, every single day. Move it to the bottom, next to the user's question. Tomorrow, check cached_tokens. The cache does not pay you for what you write. It pays you for what you refuse to change.
顯示更多
0
13
49
1
轉發到社區