
NVIDIA H100 GPU Cloud
Spin up NVIDIA H100 SXM GPUs in minutes from $2.98/hr, the proven Hopper workhorse for large-scale training, fine-tuning, and high-throughput inference.
from $2.98/hr
- GPU memory
- 80GB HBM3
- Memory bandwidth
- 3.35 TB/s
Technical specifications
- Architecture
- Hopper
- GPU memory
- 80GB HBM3
- Memory bandwidth
- 3.35 TB/s
- NVLink
- 900 GB/s
- FP16/BF16 (Tensor)
- 1,979 TFLOPS
- FP8 (Tensor)
- 3,958 TFLOPS
- Max TDP
- 700W
- GPUs per node
- 8 (HGX H100)
*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.
Pricing & availability
What's the H100 best for?
Large-scale LLM training & fine-tuning
Run 70B–400B-class training on a proven Hopper stack: FP8 tensor cores, multi-node InfiniBand, and NeMo and Megatron support.
High-throughput inference
H100 FP8 keeps token throughput high and tail latency low for production LLM serving.
Research & HPC
Mature CUDA, PyTorch, JAX, and NeMo ecosystem: spin up H100 nodes on demand for experiments, scientific compute, and deadline-driven runs.
Compare NVIDIA data-center GPUs
| A100 Ampere | H100 You're viewing | H200 Hopper | B200 Blackwell | B300 Blackwell | GB300 Blackwell | |
|---|---|---|---|---|---|---|
| Architecture | Ampere | Hopper | Hopper | Blackwell | Blackwell | Blackwell |
| GPU memory | 80GB HBM2e | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | up to 288GB HBM3e |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | 8 TB/s |
| FP8 (Tensor) | Not supported | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | 10 PFLOPS |
| Access | from $1.48/hr | from $2.98/hr | Reserved only | Reserved only | Reserved only | Reserved only |
| Best for | Cost-efficient fine-tuning & inference | Cost-efficient training & inference | Long-context & large-model inference | Frontier-scale training (FP4) | Largest models & reasoning inference | Cluster-scale training & inference |
Why industry-leading teams run GPUs on VESSL Cloud
Capacity across clouds
Access GPU capacity across clouds through one platform, without per-cloud quotas and contracts.
From one node to a cluster
Run up to 8 GPUs on one node, and move to a multi-node VM Cluster over InfiniBand when one node is not enough.
Transparent pricing
Published per-second rates for self-serve GPUs, and reserved terms quoted by GPU, term, and volume.
Enterprise-ready
SOC 2 Type II compliance, with dedicated support for production AI.
Frequently asked questions
How much does an NVIDIA H100 cost on VESSL Cloud?
H100 SXM (80GB) starts at $2.98/hr on-demand. Reserved rates are set by term and volume. Provision self-serve at cloud.vessl.ai, or reserve capacity to guarantee it.
How much VRAM does the H100 have?
The H100 SXM has 80GB of HBM3 memory at 3.35 TB/s bandwidth. If you need more memory for larger models or longer context, the H200 offers 141GB HBM3e.
What's the difference between the H100 and H200?
Both share the same Hopper compute, but the H200 carries 141GB of faster HBM3e memory (vs 80GB HBM3 on the H100) at 4.8 TB/s, fitting larger models, bigger batches, and longer context windows.
Can I run multi-node H100 training?
Yes, on a VM Cluster: HGX H100 nodes (8 GPUs each) joined by high-speed InfiniBand, with root SSH on every node and near bare-metal performance through GPU passthrough. On one node, a Workspace or Job can use up to 8 H100s.
Do you offer reserved pricing?
Reserved commitments hold your capacity, at a rate set by term and volume. Contact us.
Explore other GPUs
Different workload? Pick the GPU that fits your memory, throughput, and budget.
Same Hopper compute as the H100 with 141GB HBM3e, for long-context LLMs and larger models without sharding.
View detailsBlackwell with 192GB HBM3e and FP4 acceleration, built for frontier-scale training and high-throughput inference.
View detailsBlackwell Ultra with up to 288GB HBM3e, for the largest models and high-concurrency reasoning inference.
View details
Where AI models
get their GPUs
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support