
Per-second
billing with no long-term commitment
Hundreds of TB
of demonstrations on Object Storage
InfiniBand
multi-node VM Cluster
Trusted by teams building physical AI
Built around the robot learning loop
Run in the field, collect, curate, retrain, redeploy. In physical AI the durable advantage is the data flywheel behind the model, and most of that loop is infrastructure, not modeling. Demo success tracks how fast it turns.
Land the data next to the GPUs
Every field session adds synchronized camera streams, language instructions, and joint trajectories, and the dataset quietly grows into the hundreds of terabytes.
What slows the loop isn't capacity, it's movement: copying data into a node, preprocessing it, and pulling checkpoints back out for the robot. Storage that every workload mounts at once removes the copy-per-teammate step entirely.

$0.03
per GiB-month: episode archives live on cold-tier Object Storage
Read-write-many
concurrent workloads mount the same volume, no per-team copies
Any cluster
attach volumes to jobs at runtime, wherever the GPUs are
Sweep in bursts, pay for the burst
Recipe search is parallel by nature: augmentations, action heads, learning rates, many short runs at once. GPU demand arrives as waves, not a baseline.
On-demand GPUs billed by the second match that shape. Fan a sweep out from the CLI, release everything when the loop moves back to the lab, and switch to reserved terms once training becomes steady.

$3.19/hr
H100 SXM on-demand, self-serve
Per-second
billing with no long-term commitment
One command
vesslctl fans a sweep out across the GPUs you need
Graduate to the foundation run
When a recipe settles and the backbone goes into its big training run, a single node stops being enough, and the interconnect between nodes becomes the spec that matters.
A VM Cluster hands you the whole machine: 8 GPUs per node, from H100 to the latest Blackwell, with InfiniBand in between, root SSH, and room for the robotics toolchain (ROS 2, custom drivers, your own scheduler) that never fits a locked-down image.

8 per node
H100 to Blackwell, 1–8 nodes
InfiniBand
between nodes, on by default
Root SSH
on every node: bring ROS 2, drivers, Slurm or Kubernetes
A robot improves only as fast as its learning loop turns. Every hour the loop stalls is an hour the robot doesn't get smarter.
A workload unlike text-only training
Training models that act in the real world stresses different parts of the stack than training chatbots.
Multimodal by default
A training sample isn't a sentence. It's synchronized vision, language instructions, and robot state. Data pipelines and preprocessing carry far more of the load than in text-only training.
Real-world data, hundreds of terabytes
Demonstration data is collected on physical robots, not scraped from the web. Accumulated experiments quickly grow into hundreds of terabytes that have to move between robots and GPUs.
Simulation and RL, not just supervised passes
Much of robot learning is reinforcement learning in simulation: thousands of parallel environments, short bursty rollouts, then quiet while policies are validated on the robot. Models stay compact for on-device inference, so GPU demand arrives in waves, not a steady baseline.
Train in the cloud, act on the robot
Inference runs on the robot itself, where latency is a physical-safety problem. The cloud's job is training throughput, and fast checkpoint round-trips to get new models onto hardware.
Where GPU infrastructure gets in the way
The blockers robotics teams describe are rarely about raw GPU speed.
GPU access that can't keep up
Sweeps need eight GPUs today and two tomorrow. Waitlists, quota approvals, and allocations that fail on restart break the experiment cadence robot learning depends on.
Checkpoint and dataset round-trips
Every iteration uploads demonstration data and pulls checkpoints back down for on-robot testing. Storage and network quality, not GPU speed, often set the pace of iteration.
CPU bottlenecks nobody priced in
Image preprocessing and simulation rendering are CPU-hungry. A node with fast GPUs and starved CPUs still stalls the training loop.
Environment setup that breaks the run
Robotics stacks pin exact versions of CUDA, drivers, and tools like Isaac Sim, MuJoCo, ROS 2, and PhysX. One mismatch and the run doesn't slow down. It won't start. Teams can lose days to setup before a GPU does any useful work.
How VESSL Cloud fits
Compute plus workflow, mapped to what robot learning actually needs.
Self-serve GPUs, billed per second
Pick up GPUs when a sweep starts and release them when it ends, self-serve, with no quota tickets to file. When training becomes steady, reserved terms from three months guarantee capacity.
VM Cluster with root access
Scale to dedicated multi-node clusters: GPU nodes wired with InfiniBand, root SSH, and full driver control, as close to bare metal as a cloud gets.
Reproducible robotics environments
Pin your CUDA, simulator, and ROS 2 stack into a container image once. Every job and workspace starts from that exact environment, so a run that worked yesterday works today, and a teammate gets the same setup without rebuilding it.
Storage for robot-scale data
Object Storage holds hundreds-of-terabyte demonstration datasets at lower cost. Cluster Storage shares files across concurrent workloads with fast network performance. Attach volumes to any job.
One-command training sweeps
Define a job once with vesslctl and fan it out across the GPUs you need, with logs, metrics, and artifacts in a single view.
Workspaces for the dev loop
Connect VS Code or Cursor to a persistent GPU workspace, iterate interactively, and pause to keep the environment without paying for idle GPUs.
Physical AI on VESSL Cloud: FAQ
What GPUs do Physical AI teams train on?
Most robot foundation model training today runs on H100-class GPUs. On VESSL Cloud, H100 SXM starts at $3.19/hr self-serve, with H200 and Blackwell GPUs available on request. Inference usually runs on the robot's own edge hardware, so cloud spend stays focused on training.
How do I move hundreds of terabytes of robot data to the cloud?
Create an Object Storage volume, load data from the web console or the vesslctl CLI, and attach it to jobs at runtime. Cluster Storage covers fast file sharing between concurrent workloads.
Do I need a long-term contract to train on VESSL Cloud?
No. On-demand GPUs are billed per second with no commitment. When training becomes steady, reserved capacity starts at three-month terms with up to 15% off on-demand pricing.
Can I run Isaac Sim, MuJoCo, or ROS 2 on VESSL Cloud?
Yes. Simulators and middleware run as standard GPU jobs: pin your CUDA, simulator, and ROS 2 versions into a container image, and every job and workspace starts from that exact environment. Open VLA models such as NVIDIA's Isaac-GR00T fine-tune the same way, as multi-GPU jobs.

Where AI models
get their GPUs.
- Start in minutes
- Scale to multi-node clusters
- Capacity reserved to your timeline
- Dedicated support