注册并分享邀请链接,可获得视频播放与邀请奖励。

shmidt
@shmidtqq
predicting the future with code / ai × finance / tg: shmidtqq
加入 October 2021
899 正在关注    10.9K 粉丝
There is one line in your system prompt that multiplies your bill by 5. datetime. now() The cache reads your prompt left to right and stops at the first difference. A date on line three means the 60,000 tokens of documentation below it get recomputed. Every request. Forever. The same million input tokens: $3.00 if the start of the prompt changed $0.30 if it matched byte for byte Not a discount. Not an enterprise deal. An official line on the pricing page: Input Price (Cache Hit). One agent run, 30 steps: miss → $6.08 hit → $1.22 1,000 runs a day → $1.77M a year apart. On one line of text. Moonshot themselves run 90%+ cache hit on coding traffic. That is their operating norm, not a blog aspiration. If you are at 30%, it is not the provider. 30 seconds, right now: grep -rn "now()\|uuid4()\|random" your_prompt_builder.py Everything it finds above your documents is billed at 10x, every single day. Move it to the bottom, next to the user's question. Tomorrow, check cached_tokens. The cache does not pay you for what you write. It pays you for what you refuse to change.
显示更多
0
13
49
1
转发到社区