NVIDIA today announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production.
$NBIS was named as the first AI cloud to adopt.
Groq 3 LPX extends NVIDIA's Vera Rubin platform and is built for ultrafast token generation, targeting agentic AI workloads where latency matters.
In Artificial Analysis benchmarking, it reached a record 3,400 output tokens per second running Gemma 4 31B with a 100K-token context, while NVIDIA says it can deliver 4x faster responsiveness for agents and latency-sensitive workloads vs. the nearest alternative platform.
Nebius plans to bring Groq 3 LPX to Nebius Token Factory, giving developers access to the accelerator through its production inference platform and existing API.
显示更多