
NVIDIA A100 GPU Cloud
Spin up NVIDIA A100 SXM (80GB) GPUs in minutes from $1.39/hr, the value pick for fine-tuning, cost-efficient inference, and research on a mature Ampere stack.
from $1.39/hr
- GPU memory
- 80GB HBM2e
- Memory bandwidth
- 2.0 TB/s
Technical specifications
- Architecture
- Ampere
- GPU memory
- 80GB HBM2e
- Memory bandwidth
- 2.0 TB/s
- NVLink
- 600 GB/s
- FP16/BF16 (Tensor)
- 624 TFLOPS
- Max TDP
- 400W
- GPUs per node
- 8 (HGX A100)
*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.
Pricing & availability
What's the A100 best for?
Fine-tuning & LoRA
Fine-tune 7B–70B models cost-effectively. 80GB HBM2e fits most full and LoRA fine-tuning jobs, and per-second billing keeps each iteration cheap.
Cost-efficient inference
Serve most open models in production at a lower $/hr than Hopper, with mature TensorRT, vLLM, and Triton support.
Research & teaching
The most battle-tested data-center GPU: CUDA, PyTorch, and JAX just work. Spin up A100 nodes on demand for experiments, coursework, and reproducible runs.
Compare NVIDIA data-center GPUs
| A100 You're viewing | H100 Hopper | H200 Hopper | B200 Blackwell | B300 Blackwell | GB300 Blackwell | |
|---|---|---|---|---|---|---|
| Architecture | Ampere | Hopper | Hopper | Blackwell | Blackwell | Blackwell |
| GPU memory | 80GB HBM2e | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | up to 288GB HBM3e |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | 8 TB/s |
| FP8 (Tensor) | — | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | 10 PFLOPS |
| Access | from $1.39/hr | from $3.19/hr | from $4.29/hr | from $5.79/hr | from $7.29/hr | Reserved only |
| Best for | Cost-efficient fine-tuning & inference | Cost-efficient training & inference | Long-context & large-model inference | Frontier-scale training (FP4) | Largest models & reasoning inference | Cluster-scale training & inference |
Why industry-leading teams run GPUs on VESSL Cloud
Start in minutes
Access capacity across clouds through one platform: skip quotas and procurement.
Scale to multi-node
Spin up a single GPU or scale to large multi-node clusters over high-speed InfiniBand, as much as you need.
Transparent pricing
On-demand and reserved options with pay-as-you-go billing.
Enterprise-ready
SOC 2 Type II compliance, with dedicated support for production AI.
Frequently asked questions
How much does an NVIDIA A100 cost on VESSL Cloud?
A100 SXM (80GB) starts at $1.39/hr on-demand, self-serve at cloud.vessl.ai. Reserved commitments (3+ months) lower the rate and guarantee capacity.
How much VRAM does the A100 have?
The A100 SXM has 80GB of HBM2e memory at 2.0 TB/s bandwidth, enough for most fine-tuning and inference. For larger models, longer context, or FP8, the H100 (80GB HBM3) and H200 (141GB HBM3e) add more compute and bandwidth.
A100 vs H100: which should I pick?
The A100 is the value pick: a lower $/hr for fine-tuning, inference, and research. The H100 adds FP8 and roughly 2x the tensor throughput for large-scale training and high-throughput serving. Start on the A100 and move up when a run needs it.
Can I run multi-node A100 training?
Yes. VESSL Cloud provisions HGX A100 nodes (8 GPUs each) with high-speed InfiniBand for distributed training, plus auto-checkpointing.
Do you offer reserved pricing?
Reserved commitments (3+ months) lower the rate and guarantee capacity. Contact us.
Explore other GPUs
Different workload? Pick the GPU that fits your memory, throughput, and budget.
The proven Hopper workhorse: best price/performance for training, fine-tuning, and inference. From $3.19/hr.
View detailsSame Hopper compute as the H100 with 141GB HBM3e, for long-context LLMs and larger models without sharding.
View detailsBlackwell with 192GB HBM3e and FP4 acceleration, built for frontier-scale training and high-throughput inference.
View details
Where AI models
get their GPUs.
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support