Liwa / Insights / Vera CPU & Vera Rubin

NVIDIA's Vera CPU & Vera Rubin: the real cost, and the real timeline

Liwa Insights·6 March 2026·9 min read

At CES in January, Jensen Huang said the words every hyperscaler was waiting for: Rubin is in full production. So the next NVIDIA platform is real, it is shipping this year, and it has a name for its brain. That name is Vera. The harder question is the one the keynote skipped: what does a Vera Rubin rack actually cost to own, and what has to be true about the room it lands in?

Start with the CPU nobody is talking about

Everyone fixates on the GPU. But Rubin's host processor, Vera, is a statement in its own right: 88 custom Arm "Olympus" cores, 176 threads, up to 1.5 TB of LPDDR5X, ~1.2 TB/s of memory bandwidth, and an 1.8 TB/s NVLink-C2C link straight into the GPU. NVIDIA claims roughly 2× the performance of the Grace CPU it replaces (NVIDIA, NVIDIA Technical Blog).

Why design a bespoke CPU at all, when Intel and AMD will sell you one off the shelf? Because in a rack where the GPUs are the point, the CPU's only job is to feed them without ever becoming the bottleneck. That is a coherence-and-bandwidth problem, not a raw-cores problem. Vera is a fabric component pretending to be a processor.

If the CPU is now designed backwards, from the interconnect inward, what else in the data center are we still designing forwards, from the box outward?

The rack is the computer

The flagship configuration, the Vera Rubin NVL72, fuses 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled cabinet stitched together by NVLink 6, with around 260 TB/s of scale-up bandwidth inside the rack (VideoCardz). The unit of purchase is no longer a server. It is a rack that behaves like one enormous accelerator, and it draws power like a small substation.

That reframing matters more than any spec. You don't "install some GPUs." You energise a cabinet that wants well over a hundred kilowatts and refuses to be cooled by moving air.

The timeline, plainly

Q1 2026"Full production" declared at CES
Q3 2026First system shipments
Q4 2026Volume ramp

NVIDIA brought the schedule forward: earlier guidance pointed at a second-half 2026 mass-production window, but the CES keynote moved that to full production now, with first shipments in Q3 and the volume ramp in Q4 (Data Center Dynamics, ServeTheHome). Cloud availability is lining up behind it: AWS, Google Cloud, Microsoft and Oracle, plus NVIDIA Cloud Partners like CoreWeave, Lambda, Nebius and Nscale. Nebius has committed to NVL72 in the US and Europe from H2 2026.

Now the number that should make you sit up

NVIDIA doesn't publish rack pricing. So the market estimates. The Futurum Group pegged Vera Rubin at roughly 25% more than Grace Blackwell, call it $3.5 to $4M a rack. But the more revealing estimate breaks down the bill of materials at around $7.8M, and the reason isn't the logic die. It's memory: reporting points to a ~57% jump in GPU cost and a startling ~435% surge in memory expense, with HBM4 and LPDDR5X alone accounting for roughly $2M of that $7.8M (Wccftech, Baltimore Chronicle). These are estimates, not invoices, but the direction is unmistakable.

When memory becomes the most expensive thing in the rack, the constraint stops being "can I get a GPU?" and becomes "can anyone make enough HBM?" Who actually controls your roadmap then: the chip designer, or the memory maker?

The part the spec sheet hides

A $7.8M rack is a depreciating asset the moment it powers on. Its value is destroyed by two things: idle time, and the next generation. So the only rational way to own one is to run it hot, continuously, at the lowest possible cost per token. And the largest controllable line in that equation isn't the hardware you already bought. It's the power you'll pour through it for the next three to five years, and the cooling that decides whether it throttles.

A multi-million-dollar accelerator throttling because the room can't shed its heat is the most expensive mistake in modern infrastructure. Vera Rubin doesn't just raise the performance ceiling; it raises the stakes on the building.

Where this meets Liwa

This is precisely the problem we built Liwa around. A rack like Vera Rubin NVL72 is rationed silicon you fight to acquire. The facility it needs, though, is something you can secure today: liquid-cooled space rated to 150 kW/rack, power metered at $0.10/kWh, in a UAE free zone. You bring the silicon and your brand; we guarantee the shell that keeps it running flat-out. The chip is the lottery. The power and the cooling are the part you can actually lock in.

Questions we're sitting with

Thinking about where your Rubin racks will live?

Reserve liquid-cooled, 150 kW-ready space at $0.10/kWh. Bring your own GPUs, run under your own brand.

Sources

  1. NVIDIA, Vera CPU
  2. NVIDIA Technical Blog, Inside the Rubin platform
  3. VideoCardz, Vera Rubin NVL72 detailed
  4. Data Center Dynamics, full production at CES
  5. ServeTheHome, Rubin launch at CES 2026
  6. Wccftech, memory price surge in the Vera Rubin BOM
  7. Nebius, NVL72 in US/Europe from H2 2026

Figures are public estimates and vendor statements as of May 2026; pricing is not officially disclosed by NVIDIA and will vary by configuration and contract.