註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Simon Willison
@simonw
Creator @datasetteproj, co-creator Django. PSF board. Hangs out with @natbat. He/Him. Mastodon: Bsky:
加入 November 2006
5.7K 正在關注    203.4K 粉絲
GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point?
Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation efficiency by more than 15%.
顯示更多
0
43
365
12
轉發到社區