
NVIDIA A100 GPU Cloud
Spin up NVIDIA A100 SXM (80GB) GPUs in minutes from $1.48/hr, the value pick for fine-tuning, cost-efficient inference, and research on a mature Ampere stack.
from $1.48/hr
- GPU memory
- 80GB HBM2e
- Memory bandwidth
- 2.0 TB/s
Technical specifications
- Architecture
- Ampere
- GPU memory
- 80GB HBM2e
- Memory bandwidth
- 2.0 TB/s
- NVLink
- 600 GB/s
- FP16/BF16 (Tensor)
- 624 TFLOPS
- Max TDP
- 400W
- GPUs per node
- 8 (HGX A100)
*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.
Pricing & availability
What's the A100 best for?
Fine-tuning & LoRA
Fine-tune 7B–70B models cost-effectively. 80GB HBM2e fits most full and LoRA fine-tuning jobs, and per-second billing keeps each iteration cheap.
Cost-efficient inference
Serve most open models in production at a lower $/hr than Hopper, with mature TensorRT, vLLM, and Triton support.
Research & teaching
The most battle-tested data-center GPU: CUDA, PyTorch, and JAX just work. Spin up A100 nodes on demand for experiments, coursework, and reproducible runs.
Compare NVIDIA data-center GPUs
| A100 You're viewing | H100 Hopper | H200 Hopper | B200 Blackwell | B300 Blackwell | GB300 Blackwell | |
|---|---|---|---|---|---|---|
| Architecture | Ampere | Hopper | Hopper | Blackwell | Blackwell | Blackwell |
| GPU memory | 80GB HBM2e | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | up to 288GB HBM3e |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | 8 TB/s |
| FP8 (Tensor) | Not supported | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | 10 PFLOPS |
| Access | from $1.48/hr | from $2.98/hr | Reserved only | Reserved only | Reserved only | Reserved only |
| Best for | Cost-efficient fine-tuning & inference | Cost-efficient training & inference | Long-context & large-model inference | Frontier-scale training (FP4) | Largest models & reasoning inference | Cluster-scale training & inference |
Why industry-leading teams run GPUs on VESSL Cloud
Capacity across clouds
Access GPU capacity across clouds through one platform, without per-cloud quotas and contracts.
From one node to a cluster
Run up to 8 GPUs on one node, and move to a multi-node VM Cluster over InfiniBand when one node is not enough.
Transparent pricing
Published per-second rates for self-serve GPUs, and reserved terms quoted by GPU, term, and volume.
Enterprise-ready
SOC 2 Type II compliance, with dedicated support for production AI.
Frequently asked questions
How much does an NVIDIA A100 cost on VESSL Cloud?
A100 SXM (80GB) starts at $1.48/hr on-demand, self-serve at cloud.vessl.ai. Reserved commitments lower the rate and guarantee capacity.
How much VRAM does the A100 have?
The A100 SXM has 80GB of HBM2e memory at 2.0 TB/s bandwidth, enough for most fine-tuning and inference. For larger models, longer context, or FP8, the H100 (80GB HBM3) and H200 (141GB HBM3e) add more compute and bandwidth.
A100 vs H100: which should I pick?
The A100 is the value pick: a lower $/hr for fine-tuning, inference, and research. The H100 adds FP8 and roughly 2x the tensor throughput for large-scale training and high-throughput serving. Start on the A100 and move up when a run needs it.
Can I run multi-node A100 training?
Not across nodes. A Workspace or Job can use up to 8 A100s on one node. For distributed training over InfiniBand, use a VM Cluster, which runs H100 through B300.
Do you offer reserved pricing?
Reserved commitments lower the rate and guarantee capacity. Contact us.
Explore other GPUs
Different workload? Pick the GPU that fits your memory, throughput, and budget.
The proven Hopper workhorse: best price/performance for training, fine-tuning, and inference. From $2.98/hr.
View detailsSame Hopper compute as the H100 with 141GB HBM3e, for long-context LLMs and larger models without sharding.
View detailsBlackwell with 192GB HBM3e and FP4 acceleration, built for frontier-scale training and high-throughput inference.
View details
Where AI models
get their GPUs
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support