VESSL CloudVM Cluster

Dedicated multi-node GPU clusters

Rent a dedicated GPU cluster: H100 through the latest Blackwell, wired with InfiniBand, root on every node, reserved by the week.

Why VM Cluster

Built for large-scale distributed training.

Multi-node distributed training

InfiniBand is provisioned by default, giving you high-bandwidth, low-latency RDMA between nodes for data-parallel or model-parallel workloads.

Full root control

Swap GPU drivers, load custom kernel modules, and configure the OS however you need. Root access on every node.

Bare-metal performance, single-tenant

Each VM gets its GPUs through passthrough, so overhead stays minimal, and no other customers share your hardware. Full root access through sudo.

Bring your own stack

Install Slurm, Kubernetes, security agents, or your own deployment scripts. It's your cluster.

What you get

Every cluster ships with this.

Form factorDedicated VM with GPU passthrough
GPU8 per node (NVIDIA H100, H200, B200, B300)
Cluster sizeMulti-node over InfiniBand, scoped with sales
OS imageUbuntu 24.04 with CUDA preinstalled
NetworkingInfiniBand between nodes, dedicated public IP per node
Boot disk2 TiB per node
Shared storageOptional NFS, 1–200 TiB
AccessFull root on every node
Contract4–52 weeks, prepaid reserved

H100 through the latest Blackwell (H200, B200, B300, GB300), from a single node to hundreds or 1,000+ GPUs. Talk to sales to scope and provision your cluster.

Pricing

H100 SXM, list$3.19GPU/hr
NFS shared storage$0.20GiB/mo

A VM Cluster is a prepaid reserved contract of 4–52 weeks. Each node has 8 GPUs. Reserved contract rates depend on term and volume; talk to sales for a quote.

More cluster types, coming soon

VM Cluster gives you full control today. Managed options are coming for teams that want to run less of the infrastructure themselves.

Coming soon

Managed Kubernetes

Kubernetes pre-configured for GPU workloads, managed from setup through day-2: drivers, device plugins, and CUDA/NCCL as a validated set, with control-plane HA and node auto-repair.

Talk to sales
Coming soon

Managed Slurm

Bring your lab's Slurm workflow to cloud GPUs: sbatch, srun, and job accounting kept familiar, on a managed control plane and shared filesystem. Faulty nodes are detected and replaced automatically.

Talk to sales

Frequently Asked Questions

Which GPUs and cluster sizes are available?

H100 through the latest Blackwell (H200, B200, B300, GB300), from a single node to 1,000+ GPUs. Talk to sales to scope and provision it, and it appears in your dashboard.

Why a VM instead of bare metal?

You keep bare-metal-class performance through GPU passthrough and full root access through sudo, and VMs provision fast: most clusters are ready within a few business days of ordering. If you have a hard bare-metal requirement, talk to sales.

How does the contract work?

Prepaid reserved, 4 to 52 weeks. Node count and specs are fixed for the term; to change them, place a new order. There's no auto-scaling, and for longer terms talk to sales.

Can I use my VESSL Cloud credits?

No. A VM Cluster uses a separate prepaid invoice and can't be paid with VESSL Cloud credits.

What happens to my data at expiry?

All data is deleted immediately at expiry, with no grace period. Back up before the expiry date.

How much control do I have over the node?

Full root. Change GPU drivers and kernel modules, manage your own SSH keys, and run privileged containers (Docker --privileged, custom capabilities, raw device mounts). Install your own orchestrator if you want one.

Do you provide InfiniBand?

Yes. InfiniBand is provided by default on all nodes for multi-node communication.

Do you offer managed Kubernetes or Slurm?

Not yet, they're on the roadmap. Until then, you have root, so you can install and run Kubernetes, Slurm, or any orchestrator yourself.

How is this different from Workspace and Job?

Workspace and Job run on a single node, up to 8 GPUs. VM Cluster is built for multi-node distributed training over InfiniBand, for when one node isn't enough. All three are part of VESSL Cloud, the umbrella that spans single-node Workspaces and Jobs as well as dedicated multi-node clusters. Start with a Workspace or Job, and move up to a VM Cluster when you outgrow a single node.

Where AI models
get their GPUs.

  • Start in minutes
  • Scale to multi-node clusters
  • Capacity reserved to your timeline
  • Dedicated support