NVIDIA Rubin GPU Architecture: Advancing Agentic AI and Data Center Scalability
NVIDIA’s new “Rubin” GPU architecture marks a significant leap in scalable computing, targeting the next generation of agentic AI workloads. Designed for always-on inference tasks that require reasoning, planning, tool invocation, and sequential processing, Rubin delivers up to 10 times greater agentic throughput per unit of energy compared to its predecessor, Blackwell.
Innovative Architecture and Performance
At its core, Rubin features two reticle-limited compute dies connected via NVIDIA’s high-speed NV-HBI (NVIDIA High-Bandwidth Interface), all within a single package. This powerhouse GPU integrates 336 billion transistors, up to 224 streaming multiprocessors (SMs), 896 Tensor Cores equipped with a third-generation Transformer Engine, and 288 GB of HBM4 memory. The result is an impressive 50 PetaFLOPS of compute at sparse NVFP4 precision.
To maximize sustained throughput, Rubin organizes its resources into Graphics Processor Clusters (GPCs) with a centralized L2 cache. The GigaThread Engine orchestrates workflows and optimizes resource utilization, while MIG Control partitions allow the GPU to be split into multiple virtual GPUs. Additional features include an NVDEC block for accelerated video decoding and Confidential Computing with TEE-I/O, ensuring data protection at rest, in transit, and during use.
HBM4 Memory and High-Speed Interconnects
Rubin leverages 288 GB of HBM4 memory arranged in 12-high stacks, developed in collaboration with leading memory partners. HBM4 delivers up to 22 TB/s of peak memory bandwidth—a 2.8x increase over Blackwell and Blackwell Ultra, which top out at 8 TB/s. This bandwidth boost is enabled by HBM4’s doubled interface width compared to HBM3e.
The updated Tensor Memory Accelerator (TMA) further enhances data movement efficiency. For system-wide scaling, NVLink 6 provides 3,600 GB/s of all-to-all GPU communication bandwidth, while NVLink-C2C chip-to-chip enables CPU-GPU communication at 1,800 GB/s. Rubin also features a x16 PCIe Gen 6 host interface, supporting up to 256 GB/s of bandwidth for seamless connectivity with CPUs and external devices.
Tensor Cores, Sparsity, and Advanced Attention Mechanisms
Rubin’s third-generation Tensor Cores are engineered to address inference bottlenecks, doubling throughput per clock by processing twice the data along the K (reduction) dimension. This means matrix multiplications that required four K-loop iterations on Blackwell can now be completed in just two on Rubin—an advantage for models distributed across multiple GPUs.
For long-context attention, Rubin introduces structured 2:4 sparse compression for intermediate attention scores, reducing computational load in softmax and secondary attention matrix multiplications while maintaining dense outputs. Softmax performance is also improved, with up to 2x FP32 and 4x BF16/FP16 exponential throughput compared to Blackwell. Enhanced TMA support for mixture-of-experts (MoE) models allows a single descriptor to be shared across all experts, minimizing metadata and data movement overhead as expert counts scale.
Efficient Scale-Up Communication
Rubin introduces “counted writes” for device-initiated NVLink communication, streamlining GPU-to-GPU data transfers. Instead of traditional handshakes and synchronization barriers, the receiving GPU monitors a counter to confirm transfer completion, enabling each GPU to manage its own transfers and reducing latency. This approach keeps computation flowing without unnecessary stalls.
Additionally, Rubin implements fine-grained, tile-level kernel triggering. This allows consumer kernels to begin processing as soon as the required data is available, rather than waiting for the entire producer kernel to finish, minimizing idle time and maximizing throughput for inference workloads.
Power Optimization and Rack-Scale Design
Power efficiency is a key focus for Rubin, especially at rack scale. The “Vera Rubin” NVL72 system, which houses 72 GPUs, incorporates Intelligent Power Smoothing to maintain a fixed power envelope across all components. This technology reduces average power consumption by approximately 10% and peak power by about 20% over previous-generation techniques, even as workloads fluctuate rapidly.
At the data center level, NVIDIA DSX MaxLPS enables operators to provision up to 40% more GPUs within the same power budget. The third-generation MGX rack design features cable-free compute and switch trays, 45°C liquid cooling, hot-swappable NVLink switch trays, and open connectivity across NVLink and Spectrum-X Ethernet, supporting robust and flexible deployment in modern AI factories.
Conclusion
NVIDIA’s Rubin GPU architecture sets a new standard for agentic AI and scalable data center computing. With groundbreaking advances in memory bandwidth, interconnects, tensor processing, and power optimization, Rubin is poised to accelerate the development and deployment of complex AI workloads across industries.