NVIDIA is continuing to push throughput, and I think that makes sense. We need more tokens, and we need them cheaper. But what used to feel fast at 100 to 200 tokens per second is quickly becoming the new batch mode.
I spoke with
@swyx on
@latentspacepod about how expectations for inference speed are changing.