注册并分享邀请链接,可获得视频播放与邀请奖励。

Vaibhav (VB) Srivastav
@reach_vb
founder mode @OpenAI | ex @huggingface | F1 fan | Here for @at_sofdog’s wisdom | *opinions my own
加入 June 2017
299 正在关注    55.9K 粉丝
Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation efficiency by more than 15%.
显示更多
0
13
250
14
转发到社区