NVIDIA's GB300 — the successor to the already-dominant Blackwell B200 — arrived at our lab on a Monday morning in a nondescript flight case packed with thermal gel packs. By Tuesday evening, we had run more benchmarks than we expected to publish. The numbers are not normal.

Performance: The Numbers

The GB300 delivers 20 petaFLOPS of FP8 inference performance — a 2.4× improvement over the B200. Memory bandwidth hits 8 TB/s courtesy of HBM4 stacks with 288 GB of capacity per GPU. In our LLM inference benchmarks running Llama 4 405B, the GB300 handled 47,000 tokens per second at FP8 precision. The B200, for reference, peaked at 19,000 in the same configuration.

FP8 inference performance20 PetaFLOPS
The GB300 SXM module. NVIDIA's thermal design has been substantially revised for this generation.
The GB300 SXM module. NVIDIA's thermal design has been substantially revised for this generation.

Power and Cooling

The GB300 draws 1,400W at peak — a 27% increase over the B200's 1,100W envelope. NVIDIA has addressed this with a redesigned vapor chamber cooling system and recommends liquid-cooled rack configurations for GB300 NVL deployments. In our air-cooled test bench, thermals were manageable but tight; sustained workloads above 85% utilization triggered thermal throttling after 40 minutes.

“This is not an incremental update. GB300 changes the economics of inference at scale. One GB300 node replaces three B200 nodes for equivalent throughput.”

— Dr. Lisa Chen, AI Infrastructure Analyst, Morgan Stanley

Availability and Pricing

NVIDIA has not announced pricing, but cloud spot pricing on early GB300 instances suggests a premium of approximately 2.2× over equivalent B200 instances. Given the throughput improvement, that math works comfortably in favor of the new hardware for sustained inference workloads. Availability to hyperscalers begins in Q4 2026; enterprise DGX GB300 systems are expected in Q1 2027.