"If you knew you could get that many tokens, you would build different products."
Logan Kilpatrick (
@OfficialLoganK ,
@GoogleDeepMind) on why fast inference doesn't just make AI faster. It changes what is possible to build.
@googlegemma's Gemma 4 is now on Cerebras, running at over 1800 TPS.