Rubin CPX: NVIDIA splits inference into two kinds of chip
For years a GPU was a GPU. You bought the fastest one you could and pointed it at everything. Then NVIDIA did something quietly radical: it cut the act of inference in half and built a different chip for each half. When the company that designs the world's AI silicon decides one accelerator is no longer the right shape, the building that houses them has to change too. And one of the workloads driving the split is the one we care about most.
The split, in numbers
Announced in September 2025, the Rubin CPX is a purpose-built GPU for the context phase of inference, the compute-heavy work of ingesting a huge prompt before a single word comes out. NVIDIA's insight is that inference has two phases with opposite needs: the context phase is compute-bound, while the generation phase, producing tokens one by one, is memory-bandwidth-bound (NVIDIA, The Next Platform).
Rubin CPX brings 30 petaFLOPs of NVFP4 compute, roughly 3x the attention acceleration of a GB300 NVL72, and 128 GB of cost-efficient GDDR7. It slots into the Vera Rubin NVL144 CPX rack, which combines 144 Rubin GPUs, 144 Rubin CPX, 36 Vera CPUs and around 8 exaFLOPs of NVFP4, arriving with Vera Rubin in 2026. Its headline use cases: million-token coding and generative video (NVIDIA Technical Blog).
Why "one chip for everything" is ending
The logic is the logic of any maturing industry: specialise. A general-purpose chip is wonderful until the workload is big enough that the wasted parts start to cost real money. Long-context inference is exactly that workload. The context phase wants raw compute and cheaper memory; the generation phase wants blistering bandwidth. Forcing both onto one identical chip means overpaying for half the job, all the time.
So NVIDIA disaggregated. Different chips, matched to different phases, wired together in one rack. It's more efficient, and it's a tell about where inference is heading: toward heterogeneous racks, tuned per workload, denser and more power-hungry than the uniform server farm they replace.
The workload they named: generative video
Pay attention to the use cases NVIDIA chose to highlight. Not chatbots. Million-token coding and generative video, the most demanding long-context inference there is. Video generation devours compute on the context side and bandwidth on the generation side, which is precisely why it needs a disaggregated, ultra-dense rack rather than a commodity server. The chip was shaped around this kind of work.
That matters because generative video is not a niche. It's one of the fastest-growing inference workloads, and it runs continuously, at scale, sensitive to every cent of cost per second of output. The facilities that can host these dense, specialised racks, and power them cheaply, are the ones that will run the next wave of AI media.
Rubin CPX was built, in NVIDIA's own words, for generative video, the exact workload Liwa exists to serve. A disaggregated, 8-exaFLOP rack is dense and thirsty for power, which is why it needs a hall like Liwa's: liquid-cooled, rated to 150 kW/rack, with power locked at $0.10/kWh. As inference splits into specialised, heterogeneous racks, the winning operator is the one who can power and cool them cheaply, under their own brand. That's the whole design.
Questions we're sitting with
- If the chip designer is specialising silicon by workload phase, can a generic, air-cooled facility still host the result?
- Generative video was named as the target workload. Is your infrastructure ready for the most demanding inference there is?
- When racks become heterogeneous and denser, does the operator's edge move decisively to power and cooling?
Host the dense, specialised racks the next wave needs.
Reserve liquid-cooled, 150 kW-ready capacity at $0.10/kWh, built for generative video and long-context inference, your hardware, your brand.
Sources
- NVIDIA, Unveils Rubin CPX for massive-context inference
- NVIDIA Technical Blog, Rubin CPX for 1M+ token context
- The Next Platform, NVIDIA disaggregates long-context inference
- Tom's Hardware, Rubin CPX and the disaggregated inference architecture
Specifications are NVIDIA statements as of September 2025; Rubin CPX and the Vera Rubin NVL144 CPX rack are slated to ship with Vera Rubin in 2026 and details may change.