Finally.
Huge deals for $NBIS and great for customers.
Should be very important for customers for example in Finance and other areas where
speed = money
NVIDIA today announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production.
$NBIS was named as the first AI cloud to adopt.
Groq 3 LPX extends NVIDIA's Vera Rubin platform and is built for ultrafast token generation, targeting agentic AI workloads where latency matters.
In Artificial Analysis benchmarking, it reached a record 3,400 output tokens per second running Gemma 4 31B with a 100K-token context, while NVIDIA says it can deliver 4x faster responsiveness for agents and latency-sensitive workloads vs. the nearest alternative platform.
Nebius plans to bring Groq 3 LPX to Nebius Token Factory, giving developers access to the accelerator through its production inference platform and existing API.