Reference
Sources & Method
Three rules govern every number on this site: measured figures and announced figures are never mixed; where sources disagree the range is shown rather than a pick; and any citation whose document has not actually been read is marked as such rather than quietly padding the list. Where verifying a citation showed the underlying claim to be wrong, the claim was changed and the change is recorded below.
Figures the sources disagree about
These appear throughout the site marked ±, and hovering one shows its range. Corroborate each against a primary datasheet before treating it as settled.
Pooled HBM3e per rack ~13.5 TB
13.4 – 13.8 TB
Sources differ on rounding, on physical versus usable-after-ECC capacity, and on SKU. NVIDIA’s 192 GB per GPU gives 13.82 TB physical and 12.96 TB after ECC; Supermicro’s datasheet quotes up to 372 GB per Superchip — 186 GB per GPU — which is where the 13.4 TB figure comes from.
Copper cables in the NVLink spine >5,000
5,000 – 5,184
NVIDIA’s OCP contribution says "over 5,000"; other sources cite 5,184. Jensen Huang described it as "5,000 NVLink cables. In total, 2 miles."
Rack power ~120 kW
120 kW nominal · 125–135 kW operating (Supermicro) · 132 kW fully loaded (Schneider Electric)
Supermicro’s datasheet states an operating power of 125–135 kW and 132 kW of installed power-shelf capacity. Steven Carlini, writing for Schneider Electric: "When fully loaded into a rack, the latest NVIDIA-based GPU servers require 132 kW of power." The commonly quoted ~120 kW is the nominal design figure, not a measured ceiling.
Coolant flow ~2–3 L/min per module
2–3 L/min per module · 30–40 L/min per rack (NVIDIA ACS) · up to ~130 L/min per rack (QCT)
Figures vary by whether they are quoted per cold plate or per rack, and by the assumed ΔT. Always state the basis.
Cost of liquid cooling per MW ~$2M retrofit
~$2M per MW to retrofit · upwards of $11M per MW for a new greenfield liquid-cooled build
STL Partners, May 2026. A widely repeated "$5–10M per MW" retrofit figure is often attributed to Schneider Electric; it does not appear in the Schneider article this site cites, and no primary source for it could be found — so it is not used here.
HBM3e per Blackwell Ultra GPU 288 GB
279 GB vs 288 GB
SKU and ECC accounting. Note that NVIDIA’s own MLPerf Training v5.1 blog quotes 279 GB, so this is not simply a case of secondary sources getting it wrong.
Announced, not measured
Roadmap figures drawn from keynote material and supply-chain reporting. They are labelled at every point of use and are never placed beside benchmark results as if they were the same kind of claim.
Vera Rubin NVL144 FP4 ~3.6 EF
Announced for 2H 2026. Note that "144" counts dies, not packages — there are 72 Rubin packages.
Next-generation rack power 240 kW
Schneider Electric, forward-looking: "The next generation, expected in under a year, will require 240 kW per rack."
Known gaps
All sources
Primary
GB200 NVL72 product page NVIDIA
Source of the "acts as a single, massive GPU" framing and the 30× / 10× inference claims, with their measured configuration.
GB300 NVL72 product page NVIDIA
Grace CPU Superchip NVIDIA
Describes the standalone Grace Superchip. The GB200 Grace is configured differently — do not capacity-plan from this page.
NVLink and NVLink Switch NVIDIA
NVIDIA GB200 NVL72 Delivers Trillion-Parameter LLM Training and Real-Time Inference NVIDIA Technical Blog
Confirms 900 GB/s NVLink-C2C, 30 TB unified memory, 130 TB/s fabric, and the 30× GPT-MoE-1.8T figure.
NVIDIA Blackwell Architecture Sweeps MLPerf Training v5.1 Benchmarks NVIDIA Technical Blog
GB300 NVL72 debut in MLPerf Training; NVFP4 used in training for the first time; the 10-minute Llama 3.1 405B record on more than 5,000 Blackwell GPUs.
How NVIDIA GB200 NVL72 and NVIDIA Dynamo Boost Inference Performance for MoE Models NVIDIA Technical Blog
Scaling Large MoE Models with Wide Expert Parallelism on NVL72 Rack-Scale Systems NVIDIA Technical Blog
How New GB300 NVL72 Features Provide Steady Power for AI NVIDIA Technical Blog
The power-smoothing source: programmable power caps, integrated electrolytic capacitors, a hardware power burner, and a measured 30% reduction in peak grid demand training Megatron.
Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era NVIDIA Technical Blog
MNNVL User Guide — Overview NVIDIA Docs
NVIDIA IMEX Service for NVLink Networks — Overview NVIDIA Docs
IMEX brokers GPU memory export/import across OS domains over NVLink; it does not depend on CUDA and communicates over TCP and gRPC.
GB200 NVL Multi-Node Tuning Guide — Power and Thermals NVIDIA Docs
Names Power Smoothing as implemented for bulk synchronous workloads, and covers power balancing within a provisioned rack limit.
MLPerf Inference & Training results MLCommons
Workers Static Assets Cloudflare Docs
Workers pricing Cloudflare Docs
R2 pricing Cloudflare Docs
Independent analysis
Nvidia's Optical Boogeyman — NVL72, InfiniBand Scale Out, 800G & 1.6T Ramp SemiAnalysis
Running the NVLink spine over optics would add roughly 20 kW for transceivers and retimers alone. Also the origin of the 5,184 cable count.
This is the NVIDIA DGX GB200 NVL72 ServeTheHome
Teardown coverage: half-width nodes two-abreast in 1U, nine switch trays of two chips each, four ports and 18 links per chip, power shelves and the CDU below the compute nodes.
Global Data Center Survey 2025 — mean rack density 7.6 kW Uptime Institute
NVLink and Grace reference notes glennklockwood.com
Why Liquid Cooling For AI Data Centers Is Harder Than It Looks Steven Carlini, Schneider Electric — Forbes Technology Council
Confirmed verbatim: "When fully loaded into a rack, the latest NVIDIA-based GPU servers require 132 kW of power" and "The next generation, expected in under a year, will require 240 kW per rack." This article does NOT contain the retrofit-cost figure that is often attributed to it.
The Retrofitting Roadmap: An Evolution of Liquid Cooling STL Partners (supported by Airedale)
Liquid cooling retrofits at "around USD2 million per MW" against "upwards of USD11 million per MW" for new greenfield liquid-cooled builds.
Vendor / integrator
CoreWeave, NVIDIA and IBM Set MLPerf Record with the Largest GB200 Blackwell Cluster CoreWeave
Confirmed verbatim: "2,496 NVIDIA Blackwell GPUs across 39 racks, each containing 64 active GPUs", Llama 3.1 405B "completed in 27.33 minutes", against "around 156 racks" for an equivalent H100 setup at 32 GPUs per rack.
Supermicro NVIDIA GB200 NVL72 SuperCluster datasheet (PDF) Supermicro
Read directly. Rack 2236 × 600 × 1068 mm; 8 × 1U 33 kW power shelves totalling 132 kW; operating power 125–135 kW; 10 + 8 compute trays around 9 NVLink switch trays; up to 372 GB HBM3e and 480 GB LPDDR5X per Superchip; in-rack 250 kW CDU.
QCT QoolRack Stand-Alone — Advanced Liquid Cooling for NVIDIA GB200 NVL72 Systems (PDF) QCT unconfirmed
The document exists and resolves, but its text is embedded as CID-encoded fonts and could not be extracted, so the 45 °C maximum inlet, 65 °C maximum return and ~130 L/min per-rack figures attributed to it are still second-hand here. Several secondary sources repeat exactly these numbers and credit QCT. Confirm against the readable document before treating them as primary.
Standards contribution
NVIDIA Contributes NVIDIA GB200 NVL72 Designs to the Open Compute Project NVIDIA Technical Blog
Specifies four NVLink cartridges with over 5,000 copper cables delivering 260 TB/s AllReduce bandwidth, a 1,400 A busbar, and over 100 lb of rack reinforcement steel.
Pedagogy reference
Explorable explainers (Gears, Cameras and Lenses, Internal Combustion Engine) Bartosz Ciechanowski
Explorable Explanations Bret Victor
Method
The fact base is a single module — every figure the site states is defined once, with its unit, its sources, any disputed range, and whether it is announced. Pages reference figures by identifier rather than typing numbers into prose, so a correction lands everywhere at once and it is not possible for two chapters to disagree with each other.
The 3D rack is original geometry generated procedurally from those same dimensions. NVIDIA publishes no CAD for this rack, so an original model was required regardless; generating it from parameters rather than authoring a mesh keeps the dimensions next to the sourced figures they come from, ships no asset bytes, and leaves every part individually addressable for the exploded view. It is a schematic diagram in three dimensions, not a fabrication drawing.
No NVIDIA artwork, branding or CAD is reproduced anywhere on this site. Word marks are used nominatively to describe the hardware. Mechanical proportions follow public teardown photography and openly licensed Open Compute Project documentation.