@cerebras is now running
@GoogleDeepMind's Gemma 4 - the leading open-weight multimodal model - at 1,851 tokens per second in public preview.
This is 35x faster than a typical GPU endpoint.
Cerebras speed also translates into world class latency - Gemma 4 on Cerebras returns its first answer token inclusive of reasoning in 1.5 seconds, making Cerebras the only provider that lets Gemma 4 be used in real-time settings.
This is the power of wafer scale.