NVIDIA's Vera CPU & Vera Rubin: the real cost, and the real timeline
At CES in January, Jensen Huang said the words every hyperscaler was waiting for: Rubin is in full production. So the next NVIDIA platform is real, it is shipping this year, and it has a name for its brain. That name is Vera. The harder question is the one the keynote skipped: what does a Vera Rubin rack actually cost to own, and what has to be true about the room it lands in?
Start with the CPU nobody is talking about
Everyone fixates on the GPU. But Rubin's host processor, Vera, is a statement in its own right: 88 custom Arm "Olympus" cores, 176 threads, up to 1.5 TB of LPDDR5X, ~1.2 TB/s of memory bandwidth, and an 1.8 TB/s NVLink-C2C link straight into the GPU. NVIDIA claims roughly 2× the performance of the Grace CPU it replaces (NVIDIA, NVIDIA Technical Blog).
Why design a bespoke CPU at all, when Intel and AMD will sell you one off the shelf? Because in a rack where the GPUs are the point, the CPU's only job is to feed them without ever becoming the bottleneck. That is a coherence-and-bandwidth problem, not a raw-cores problem. Vera is a fabric component pretending to be a processor.
The rack is the computer
The flagship configuration, the Vera Rubin NVL72, fuses 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled cabinet stitched together by NVLink 6, with around 260 TB/s of scale-up bandwidth inside the rack (VideoCardz). The unit of purchase is no longer a server. It is a rack that behaves like one enormous accelerator, and it draws power like a small substation.
That reframing matters more than any spec. You don't "install some GPUs." You energise a cabinet that wants well over a hundred kilowatts and refuses to be cooled by moving air.
The timeline, plainly
NVIDIA brought the schedule forward: earlier guidance pointed at a second-half 2026 mass-production window, but the CES keynote moved that to full production now, with first shipments in Q3 and the volume ramp in Q4 (Data Center Dynamics, ServeTheHome). Cloud availability is lining up behind it: AWS, Google Cloud, Microsoft and Oracle, plus NVIDIA Cloud Partners like CoreWeave, Lambda, Nebius and Nscale. Nebius has committed to NVL72 in the US and Europe from H2 2026.
Now the number that should make you sit up
NVIDIA doesn't publish rack pricing. So the market estimates. The Futurum Group pegged Vera Rubin at roughly 25% more than Grace Blackwell, call it $3.5 to $4M a rack. But the more revealing estimate breaks down the bill of materials at around $7.8M, and the reason isn't the logic die. It's memory: reporting points to a ~57% jump in GPU cost and a startling ~435% surge in memory expense, with HBM4 and LPDDR5X alone accounting for roughly $2M of that $7.8M (Wccftech, Baltimore Chronicle). These are estimates, not invoices, but the direction is unmistakable.
The part the spec sheet hides
A $7.8M rack is a depreciating asset the moment it powers on. Its value is destroyed by two things: idle time, and the next generation. So the only rational way to own one is to run it hot, continuously, at the lowest possible cost per token. And the largest controllable line in that equation isn't the hardware you already bought. It's the power you'll pour through it for the next three to five years, and the cooling that decides whether it throttles.
A multi-million-dollar accelerator throttling because the room can't shed its heat is the most expensive mistake in modern infrastructure. Vera Rubin doesn't just raise the performance ceiling; it raises the stakes on the building.
This is precisely the problem we built Liwa around. A rack like Vera Rubin NVL72 is rationed silicon you fight to acquire. The facility it needs, though, is something you can secure today: liquid-cooled space rated to 150 kW/rack, power metered at $0.10/kWh, in a UAE free zone. You bring the silicon and your brand; we guarantee the shell that keeps it running flat-out. The chip is the lottery. The power and the cooling are the part you can actually lock in.
Questions we're sitting with
- If memory is now the gating cost, does the "GPU shortage" quietly become an HBM shortage, and does that change who you should be negotiating with?
- At $7.8M a rack, what utilisation rate turns it from a trophy into a return? And what power price makes that math survive a downturn?
- Full production in Q1, shipments in Q3: if the silicon is finally moving, is the new bottleneck simply somewhere to plug it in?
Thinking about where your Rubin racks will live?
Reserve liquid-cooled, 150 kW-ready space at $0.10/kWh. Bring your own GPUs, run under your own brand.
Sources
- NVIDIA, Vera CPU
- NVIDIA Technical Blog, Inside the Rubin platform
- VideoCardz, Vera Rubin NVL72 detailed
- Data Center Dynamics, full production at CES
- ServeTheHome, Rubin launch at CES 2026
- Wccftech, memory price surge in the Vera Rubin BOM
- Nebius, NVL72 in US/Europe from H2 2026
Figures are public estimates and vendor statements as of May 2026; pricing is not officially disclosed by NVIDIA and will vary by configuration and contract.