NVIDIA Ampere

NVIDIA A100 GPU Cloud

Spin up NVIDIA A100 SXM (80GB) GPUs in minutes from $1.39/hr, the value pick for fine-tuning, cost-efficient inference, and research on a mature Ampere stack.

from $1.39/hr

NVIDIA A100 SXM
GPU memory
80GB HBM2e
Memory bandwidth
2.0 TB/s

Technical specifications

Architecture
Ampere
GPU memory
80GB HBM2e
Memory bandwidth
2.0 TB/s
NVLink
600 GB/s
FP16/BF16 (Tensor)
624 TFLOPS
Max TDP
400W
GPUs per node
8 (HGX A100)

*Peak performance with sparsity, per NVIDIA official specs. Final specs may vary by node configuration.

Pricing & availability

NVIDIA A100 SXMSelf-serve
$1.39per GPU/hr
Start now

What's the A100 best for?

Fine-tuning & LoRA

Fine-tune 7B–70B models cost-effectively. 80GB HBM2e fits most full and LoRA fine-tuning jobs, and per-second billing keeps each iteration cheap.

Cost-efficient inference

Serve most open models in production at a lower $/hr than Hopper, with mature TensorRT, vLLM, and Triton support.

Research & teaching

The most battle-tested data-center GPU: CUDA, PyTorch, and JAX just work. Spin up A100 nodes on demand for experiments, coursework, and reproducible runs.

Compare NVIDIA data-center GPUs

A100
You're viewing
H100
Hopper
H200
Hopper
B200
Blackwell
B300
Blackwell
GB300
Blackwell
ArchitectureAmpereHopperHopperBlackwellBlackwellBlackwell
GPU memory80GB HBM2e80GB HBM3141GB HBM3e192GB HBM3eup to 288GB HBM3eup to 288GB HBM3e
Memory bandwidth2.0 TB/s3.35 TB/s4.8 TB/s8 TB/s8 TB/s8 TB/s
FP8 (Tensor)3,958 TFLOPS3,958 TFLOPS9 PFLOPS10 PFLOPS10 PFLOPS
Accessfrom $1.39/hrfrom $3.19/hrfrom $4.29/hrfrom $5.79/hrfrom $7.29/hrReserved only
Best forCost-efficient fine-tuning & inferenceCost-efficient training & inferenceLong-context & large-model inferenceFrontier-scale training (FP4)Largest models & reasoning inferenceCluster-scale training & inference

Why industry-leading teams run GPUs on VESSL Cloud

Start in minutes

Access capacity across clouds through one platform: skip quotas and procurement.

Scale to multi-node

Spin up a single GPU or scale to large multi-node clusters over high-speed InfiniBand, as much as you need.

Transparent pricing

On-demand and reserved options with pay-as-you-go billing.

Enterprise-ready

SOC 2 Type II compliance, with dedicated support for production AI.

Frequently asked questions

How much does an NVIDIA A100 cost on VESSL Cloud?

A100 SXM (80GB) starts at $1.39/hr on-demand, self-serve at cloud.vessl.ai. Reserved commitments (3+ months) lower the rate and guarantee capacity.

How much VRAM does the A100 have?

The A100 SXM has 80GB of HBM2e memory at 2.0 TB/s bandwidth, enough for most fine-tuning and inference. For larger models, longer context, or FP8, the H100 (80GB HBM3) and H200 (141GB HBM3e) add more compute and bandwidth.

A100 vs H100: which should I pick?

The A100 is the value pick: a lower $/hr for fine-tuning, inference, and research. The H100 adds FP8 and roughly 2x the tensor throughput for large-scale training and high-throughput serving. Start on the A100 and move up when a run needs it.

Can I run multi-node A100 training?

Yes. VESSL Cloud provisions HGX A100 nodes (8 GPUs each) with high-speed InfiniBand for distributed training, plus auto-checkpointing.

Do you offer reserved pricing?

Reserved commitments (3+ months) lower the rate and guarantee capacity. Contact us.

Where AI models
get their GPUs.

  • Start in minutes
  • Scale to multi-node clusters
  • Capacity reserved to your timeline
  • Dedicated support