Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model.
The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation efficiency by more than 15%.