From Blackwell to Rubin: Reading NVIDIA's Roadmap for What Comes Next

NVIDIA's data-centre roadmap moves fast, and the naming can obscure more than it clarifies. Blackwell, Blackwell Ultra, Vera Rubin — what actually changes between generations, and what should teams plan for?


Blackwell: the rack becomes the unit

The original Blackwell platform (GB200 NVL72) marked the shift from server-scale to rack-scale computing. Seventy-two GPUs and thirty-six Grace CPUs were fused into one NVLink domain, with 192 GB of HBM3e per GPU. It established the template that the rest of the roadmap follows.

Blackwell Ultra: memory and reasoning

Blackwell Ultra (GB300 NVL72) keeps the rack-scale architecture but pushes memory to 288 GB of HBM3e per GPU and adds hardware acceleration for attention, tuned specifically for AI reasoning. This is the current flagship — the part that delivers the headline 10× inference gains over Hopper.

Vera Rubin: the next platform

Vera Rubin is the generation after Blackwell Ultra. It moves to HBM4 memory (288 GB per GPU) with roughly 22 TB/s of bandwidth, introduces sixth-generation NVLink, and targets 50 PFLOPS of FP4 inference per GPU. It is the platform NVIDIA is positioning for the next wave of agentic AI at scale.

What this means for your planning

The bottom line: The roadmap points in one consistent direction: more memory, faster interconnect, and rack-scale systems tuned for inference and reasoning. Teams that plan their infrastructure around that trajectory — rather than around any single GPU — will be the ones that keep shipping.

Back to all posts Want to see where your workload sits on the roadmap? Plan ahead with us