
NVIDIA GB200 NVL72
A purpose-built rackmount platform engineered for mission-critical workloads. Dense compute, abundant memory bandwidth, and a flexible PCIe Gen5 fabric make it ideal for virtualization, databases, AI inference, and hyper-converged infrastructure.
Availability subject to compliance verification.
Grace Blackwell Superchip Overview
The Grace Blackwell Superchip represents a monumental leap in computing technology, merging unparalleled processing power with advanced architectural design. With a memory clock speed of 8Gbps provided by HBM3E technology and an expansive memory bus width of 2x2x4096-bit, this superchip achieves an astonishing memory bandwidth of 2x8TB/sec, supported by a vast 384GB of VRAM. At the core of its computational prowess, the superchip boasts tensor processing capabilities across a wide range of precisions, delivering up to 20 PFLOPS for FP4 Dense Tensor operations, 10 P(FL)OPS for INT8/FP8, 5 PFLOPS for FP16, 2.5 PFLOPS for TF32, and an impressive 90 TFLOPS for FP64 Dense Tensor calculations. Connectivity is no less advanced, with dual NVLink 5 interfaces reaching 1800GB/sec and PCIe 6.0 connections providing an additional 256GB/sec bandwidth. Powered by two "Blackwell GPUs" and harboring a staggering 416 billion transistors, this superchip is not just a powerhouse but a marvel of modern engineering. Its thermal design power (TDP) stands at 2700W, indicative of its high performance and energy demands. Fabricated on TSMC's 4NP process, the Grace Blackwell Superchip sets a new standard for high-performance computing platforms, blending the Grace and Blackwell architectures to achieve unmatched computational efficiency and throughput.

| Specification | GB200 NVL72 | GB200 Grace Blackwell Superchip |
|---|---|---|
| Configuration | 36 Grace CPU : 72 Blackwell GPUs | 1 Grace CPU : 2 Blackwell GPU |
| FP4 Tensor Core¹ | 1,440 PFLOPS | 40 PFLOPS |
| FP8/FP6 Tensor Core¹ | 720 PFLOPS | 20 PFLOPS |
| INT8 Tensor Core¹ | 720 POPS | 20 POPS |
| FP16/BF16 Tensor Core¹ | 360 PFLOPS | 10 PFLOPS |
| TF32 Tensor Core | 180 PFLOPS | 5 PFLOPS |
| FP32 | 5,760 TFLOPS | 160 TFLOPS |
| FP64 | 2,880 TFLOPS | 80 TFLOPS |
| FP64 Tensor Core | 2,880 TFLOPS | 80 TFLOPS |
| GPU Memory | Bandwidth | Up to 13.4 TB HBM3e | 576 TB/s | Up to 372 GB HBM3e | 16 TB/s |
| NVLink Bandwidth | 130 TB/s | 3.6 TB/s |
| CPU Core Count | 2,592 Arm Neoverse V2 cores | 72 Arm Neoverse V2 cores |
| CPU Memory | Bandwidth | Up to 17 TB LPDDR5X | Up to 18.4 TB/s | Up to 480 GB LPDDR5X | Up to 512 GB/s |
¹ With sparsity.