Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
NVIDIA announced that its Groq 3 LPX inference accelerator is in full production, delivering 3,400 output tokens per second on Gemma 4 31B for 100,000-token long-context use cases, 4x faster than alternatives. The system is part of the Vera Rubin platform and is being adopted by partners like SpaceXAI, CoreWeave, and Nebius.
From the source
Announced today , the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production.
blogs.nvidia.com