NVIDIA Ampere

NVIDIA A100 GPU Cloud

Spin up NVIDIA A100 SXM (80GB) GPUs in minutes from $1.48/hr, the value pick for fine-tuning, cost-efficient inference, and research on a mature Ampere stack.

from $1.48/hr

NVIDIA A100 SXM
GPU memory
80GB HBM2e
Memory bandwidth
2.0 TB/s

Technical specifications

Architecture
Ampere
GPU memory
80GB HBM2e
Memory bandwidth
2.0 TB/s
NVLink
600 GB/s
FP16/BF16 (Tensor)
624 TFLOPS
Max TDP
400W
GPUs per node
8 (HGX A100)

*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.

Pricing & availability

NVIDIA A100 SXMSelf-serve
$1.48per GPU/hr
Start now

What's the A100 best for?

Fine-tuning & LoRA

Fine-tune 7B–70B models cost-effectively. 80GB HBM2e fits most full and LoRA fine-tuning jobs, and per-second billing keeps each iteration cheap.

Cost-efficient inference

Serve most open models in production at a lower $/hr than Hopper, with mature TensorRT, vLLM, and Triton support.

Research & teaching

The most battle-tested data-center GPU: CUDA, PyTorch, and JAX just work. Spin up A100 nodes on demand for experiments, coursework, and reproducible runs.

Compare NVIDIA data-center GPUs

A100
You're viewing
H100
Hopper
H200
Hopper
B200
Blackwell
B300
Blackwell
GB300
Blackwell
ArchitectureAmpereHopperHopperBlackwellBlackwellBlackwell
GPU memory80GB HBM2e80GB HBM3141GB HBM3e192GB HBM3eup to 288GB HBM3eup to 288GB HBM3e
Memory bandwidth2.0 TB/s3.35 TB/s4.8 TB/s8 TB/s8 TB/s8 TB/s
FP8 (Tensor)Not supported3,958 TFLOPS3,958 TFLOPS9 PFLOPS10 PFLOPS10 PFLOPS
Accessfrom $1.48/hrfrom $2.98/hrReserved onlyReserved onlyReserved onlyReserved only
Best forCost-efficient fine-tuning & inferenceCost-efficient training & inferenceLong-context & large-model inferenceFrontier-scale training (FP4)Largest models & reasoning inferenceCluster-scale training & inference

Why industry-leading teams run GPUs on VESSL Cloud

Capacity across clouds

Access GPU capacity across clouds through one platform, without per-cloud quotas and contracts.

From one node to a cluster

Run up to 8 GPUs on one node, and move to a multi-node VM Cluster over InfiniBand when one node is not enough.

Transparent pricing

Published per-second rates for self-serve GPUs, and reserved terms quoted by GPU, term, and volume.

Enterprise-ready

SOC 2 Type II compliance, with dedicated support for production AI.

Frequently asked questions

How much does an NVIDIA A100 cost on VESSL Cloud?

A100 SXM (80GB) starts at $1.48/hr on-demand, self-serve at cloud.vessl.ai. Reserved commitments lower the rate and guarantee capacity.

How much VRAM does the A100 have?

The A100 SXM has 80GB of HBM2e memory at 2.0 TB/s bandwidth, enough for most fine-tuning and inference. For larger models, longer context, or FP8, the H100 (80GB HBM3) and H200 (141GB HBM3e) add more compute and bandwidth.

A100 vs H100: which should I pick?

The A100 is the value pick: a lower $/hr for fine-tuning, inference, and research. The H100 adds FP8 and roughly 2x the tensor throughput for large-scale training and high-throughput serving. Start on the A100 and move up when a run needs it.

Can I run multi-node A100 training?

Not across nodes. A Workspace or Job can use up to 8 A100s on one node. For distributed training over InfiniBand, use a VM Cluster, which runs H100 through B300.

Do you offer reserved pricing?

Reserved commitments lower the rate and guarantee capacity. Contact us.

Where AI models
get their GPUs

  • Start in minutes
  • Scale to multi-node clusters
  • Capacity reserved to your timeline
  • Dedicated support