We beat out
@nvidia's KV cache transfer for KL divergence, and we're using it to build the fastest inference
@TryTrustAI at
@ycombinator.
17.8% less TTFT, $641 less per 1M requests, and 82.5% held-out top-1 agreement on Qwen 32B, prefilled from Qwen 8B.