
Dedicated multi-node GPU clusters
Rent a dedicated GPU cluster: H100 through the latest Blackwell, wired with InfiniBand, root on every node, reserved by the week.
Why VM Cluster
Built for large-scale distributed training.

Multi-node distributed training
InfiniBand is provisioned by default, giving you high-bandwidth, low-latency RDMA between nodes for data-parallel or model-parallel workloads.

Full root control
Swap GPU drivers, load custom kernel modules, and configure the OS however you need. Root access on every node.

Bare-metal performance, single-tenant
Each VM gets its GPUs through passthrough, so overhead stays minimal, and no other customers share your hardware. Full root access through sudo.

Bring your own stack
Install Slurm, Kubernetes, security agents, or your own deployment scripts. It's your cluster.
What you get
Every cluster ships with this.
H100 through the latest Blackwell (H200, B200, B300, GB300), from a single node to hundreds or 1,000+ GPUs. Talk to sales to scope and provision your cluster.
Pricing
A VM Cluster is a prepaid reserved contract of 4–52 weeks. Each node has 8 GPUs. Reserved contract rates depend on term and volume; talk to sales for a quote.
More cluster types, coming soon
VM Cluster gives you full control today. Managed options are coming for teams that want to run less of the infrastructure themselves.

Managed Kubernetes
Kubernetes pre-configured for GPU workloads, managed from setup through day-2: drivers, device plugins, and CUDA/NCCL as a validated set, with control-plane HA and node auto-repair.
Talk to salesManaged Slurm
Bring your lab's Slurm workflow to cloud GPUs: sbatch, srun, and job accounting kept familiar, on a managed control plane and shared filesystem. Faulty nodes are detected and replaced automatically.
Talk to salesFrequently Asked Questions
Which GPUs and cluster sizes are available?
H100 through the latest Blackwell (H200, B200, B300, GB300), from a single node to 1,000+ GPUs. Talk to sales to scope and provision it, and it appears in your dashboard.
Why a VM instead of bare metal?
You keep bare-metal-class performance through GPU passthrough and full root access through sudo, and VMs provision fast: most clusters are ready within a few business days of ordering. If you have a hard bare-metal requirement, talk to sales.
How does the contract work?
Prepaid reserved, 4 to 52 weeks. Node count and specs are fixed for the term; to change them, place a new order. There's no auto-scaling, and for longer terms talk to sales.
Can I use my VESSL Cloud credits?
No. A VM Cluster uses a separate prepaid invoice and can't be paid with VESSL Cloud credits.
What happens to my data at expiry?
All data is deleted immediately at expiry, with no grace period. Back up before the expiry date.
How much control do I have over the node?
Full root. Change GPU drivers and kernel modules, manage your own SSH keys, and run privileged containers (Docker --privileged, custom capabilities, raw device mounts). Install your own orchestrator if you want one.
Do you provide InfiniBand?
Yes. InfiniBand is provided by default on all nodes for multi-node communication.
Do you offer managed Kubernetes or Slurm?
Not yet, they're on the roadmap. Until then, you have root, so you can install and run Kubernetes, Slurm, or any orchestrator yourself.
How is this different from Workspace and Job?
Workspace and Job run on a single node, up to 8 GPUs. VM Cluster is built for multi-node distributed training over InfiniBand, for when one node isn't enough. All three are part of VESSL Cloud, the umbrella that spans single-node Workspaces and Jobs as well as dedicated multi-node clusters. Start with a Workspace or Job, and move up to a VM Cluster when you outgrow a single node.

Where AI models
get their GPUs.
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support