
NVIDIA GB300 GPU Cloud
Reserve NVIDIA GB300 NVL72 capacity on VESSL Cloud: 72 Blackwell Ultra GPUs and 36 Grace CPUs per rack, unified by a 130 TB/s NVLink domain for cluster-scale training and high-concurrency reasoning inference.
- GPU memory
- up to 288GB HBM3e
- Memory bandwidth
- 8 TB/s
Technical specifications
- Architecture
- Blackwell
- GPU memory
- up to 288GB HBM3e
- Memory bandwidth
- 8 TB/s
- NVLink
- 1.8 TB/s
- FP8 (Tensor)
- 10 PFLOPS
- FP4 (Tensor)
- 20 PFLOPS
- Max TDP
- 1,400W
- GPUs per node
- 72 + 36 Grace CPUs (NVL72 rack)
*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.
Pricing & availability
What's the GB300 best for?
Cluster-scale training
A full NVL72 rack behaves like one giant accelerator: roughly 20 TB of HBM3e and a 130 TB/s NVLink domain keep trillion-parameter runs resident without hopping across slower networks.
Rack-scale reasoning inference
The unified NVLink domain spreads massive KV caches across 72 GPUs, sustaining high-concurrency agentic and long-context serving at a lower cost per token.
CPU-heavy AI pipelines
36 Grace CPUs sit on the same NVLink fabric as the GPUs, so data preprocessing and inference-time orchestration feed accelerators without PCIe bottlenecks.
Compare NVIDIA data-center GPUs
| A100 Ampere | H100 Hopper | H200 Hopper | B200 Blackwell | B300 Blackwell | GB300 You're viewing | |
|---|---|---|---|---|---|---|
| Architecture | Ampere | Hopper | Hopper | Blackwell | Blackwell | Blackwell |
| GPU memory | 80GB HBM2e | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | up to 288GB HBM3e |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | 8 TB/s |
| FP8 (Tensor) | — | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | 10 PFLOPS |
| Access | from $1.39/hr | from $3.19/hr | from $4.29/hr | from $5.79/hr | from $7.29/hr | Reserved only |
| Best for | Cost-efficient fine-tuning & inference | Cost-efficient training & inference | Long-context & large-model inference | Frontier-scale training (FP4) | Largest models & reasoning inference | Cluster-scale training & inference |
Why industry-leading teams run GPUs on VESSL Cloud
Start in minutes
Access capacity across clouds through one platform: skip quotas and procurement.
Scale to multi-node
Spin up a single GPU or scale to large multi-node clusters over high-speed InfiniBand, as much as you need.
Transparent pricing
On-demand and reserved options with pay-as-you-go billing.
Enterprise-ready
SOC 2 Type II compliance, with dedicated support for production AI.
Frequently asked questions
How do I get access to NVIDIA GB300 GPUs?
GB300 capacity is reserved and allocated on request. Talk to our team to scope an NVL72 configuration and reserve capacity ahead of your schedule.
What is the GB300 NVL72?
A rack-scale system pairing 72 Blackwell Ultra GPUs with 36 Grace CPUs in a single 130 TB/s NVLink domain, so the whole rack behaves like one large accelerator.
What's the difference between the B300 and GB300?
Both use the Blackwell Ultra GPU. The B300 ships as 8-GPU HGX nodes; the GB300 packages 72 GPUs plus 36 Grace CPUs into an NVL72 rack whose 130 TB/s NVLink domain is built for cluster-scale workloads.
How much memory does the GB300 have?
Up to 288GB HBM3e per GPU, roughly 20 TB per NVL72 rack, plus 17 TB of LPDDR5X on the Grace CPUs.
Can I reserve a full GB300 cluster?
Yes. GB300 is provisioned by the NVL72 rack, from a single rack to multi-rack deployments. Talk to our team about capacity and timelines.
Explore other GPUs
Different workload? Pick the GPU that fits your memory, throughput, and budget.
Blackwell Ultra with up to 288GB HBM3e, for the largest models and high-concurrency reasoning inference.
View detailsBlackwell with 192GB HBM3e and FP4 acceleration, built for frontier-scale training and high-throughput inference.
View detailsSame Hopper compute as the H100 with 141GB HBM3e, for long-context LLMs and larger models without sharding.
View details
Where AI models
get their GPUs.
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support