Seven chips, five racks, sold as one system
NVIDIA's announcement lists seven chips in full production: the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch, and a newly integrated Groq 3 LPU. The flagship configuration is the Vera Rubin NVL72 rack, pairing 72 Rubin GPUs with 36 Vera CPUs, alongside separate racks built around the Vera CPU and the Groq 3 chip. [1]
NVIDIA's stated performance claims for the NVL72 rack are a mixture-of-experts model trained with one-fourth the GPUs Blackwell needed, and up to ten times higher inference throughput per watt at one-tenth the cost per token. A separate 256-chip Vera CPU rack is claimed at 50 percent faster and twice the energy efficiency of a traditional CPU rack. [1]
Atlas interpretation: None of the four numbers, the GPU-count reduction, the throughput-per-watt figure, the cost-per-token figure, or the CPU efficiency claim, comes with a published benchmark methodology in NVIDIA's own release. They are vendor-reported comparisons against NVIDIA's own prior generation, not independently measured results, and neither source gathered here cites an outside benchmark that tests them. [1][2]
The Groq LPU that NVIDIA now sells
The Groq 3 LPU is not a NVIDIA design. Three months earlier, NVIDIA licensed Groq's inference technology and hired Groq's founder and much of its chip team, while Groq itself kept operating as an independent company under a non-exclusive license. The Vera Rubin lineup is the first NVIDIA platform to ship a Groq-derived chip. [1]
NVIDIA's own Groq 3 LPX rack packs 256 of the processors, each with 128GB of on-chip memory and 640 TB/s of scale-up bandwidth, and claims up to 35 times higher inference throughput per megawatt than the racks it displaces. Datacenter Knowledge reported that combining Vera Rubin GPU racks with Groq 3 LPX inference racks was pitched as a 350-fold jump in token generation, from roughly 2 million to 700 million tokens per second on a comparable system. [1][2]
Atlas interpretation: The December licensing deal read at the time as NVIDIA buying its way past a rival inference-chip maker without a full acquisition. Putting a Groq-derived chip inside its flagship platform three months later, rather than treating the license as a defensive patent play, shows the deal was for the chip itself, not just to take a competitor off the market. [1]
Doubling the forecast, on the strength of inference
Datacenter Knowledge reported that Huang, having forecast $500 billion in NVIDIA sales through 2026 a year earlier on the strength of Blackwell and Rubin GPUs, raised that projection to $1 trillion through 2027 at this GTC keynote, attributing the increase to inference demand rather than training demand. Huang framed the shift directly: AI models can now do productive work, which he called the arrival of the inflection point of inference. [2]
Atlas interpretation: The reported figure is a company sales forecast, NVIDIA's own projection of future revenue, rather than a disclosed order book or signed contracts. Neither source gathered here shows NVIDIA disclosing binding commitments behind the doubled number. [2]
The announcement carried supporting quotes from two of NVIDIA's largest customers. OpenAI's Sam Altman said NVIDIA infrastructure is the foundation that lets OpenAI keep pushing the frontier of AI, and Anthropic's Dario Amodei said the Vera Rubin platform gives Anthropic the compute, networking and system design to keep delivering. [1]
Sources
- NVIDIA Vera Rubin Opens Agentic AI Frontier
NVIDIA · Mar 16, 2026
- GTC 2026: Nvidia Unveils Vera Rubin AI Platform, Eyes $1T by 2027
Data Center Knowledge · Mar 17, 2026