VESSL CloudUse case

Training robot foundation models

Per-second

billing with no long-term commitment

Hundreds of TB

of demonstrations on Object Storage

InfiniBand

multi-node VM Cluster

Trusted by teams building physical AI

Tommoro RoboticsDiden RoboticsNdotLight

Built around the robot learning loop

Run in the field, collect, curate, retrain, redeploy. In physical AI the durable advantage is the data flywheel behind the model, and most of that loop is infrastructure, not modeling. Demo success tracks how fast it turns.

Land the data next to the GPUs

Every field session adds synchronized camera streams, language instructions, and joint trajectories, and the dataset quietly grows into the hundreds of terabytes.

What slows the loop isn't capacity, it's movement: copying data into a node, preprocessing it, and pulling checkpoints back out for the robot. Storage that every workload mounts at once removes the copy-per-teammate step entirely.

$0.03

per GiB-month: episode archives live on cold-tier Object Storage

Read-write-many

concurrent workloads mount the same volume, no per-team copies

Any cluster

attach volumes to jobs at runtime, wherever the GPUs are

Sweep in bursts, pay for the burst

Recipe search is parallel by nature: augmentations, action heads, learning rates, many short runs at once. GPU demand arrives as waves, not a baseline.

On-demand GPUs billed by the second match that shape. Fan a sweep out from the CLI, release everything when the loop moves back to the lab, and switch to reserved terms once training becomes steady.

$3.19/hr

H100 SXM on-demand, self-serve

Per-second

billing with no long-term commitment

One command

vesslctl fans a sweep out across the GPUs you need

Graduate to the foundation run

When a recipe settles and the backbone goes into its big training run, a single node stops being enough, and the interconnect between nodes becomes the spec that matters.

A VM Cluster hands you the whole machine: 8 GPUs per node, from H100 to the latest Blackwell, with InfiniBand in between, root SSH, and room for the robotics toolchain (ROS 2, custom drivers, your own scheduler) that never fits a locked-down image.

8 per node

H100 to Blackwell, 1–8 nodes

InfiniBand

between nodes, on by default

Root SSH

on every node: bring ROS 2, drivers, Slurm or Kubernetes

A robot improves only as fast as its learning loop turns. Every hour the loop stalls is an hour the robot doesn't get smarter.

The pattern behind every Physical AI conversation we have

A workload unlike text-only training

Training models that act in the real world stresses different parts of the stack than training chatbots.

Multimodal by default

A training sample isn't a sentence. It's synchronized vision, language instructions, and robot state. Data pipelines and preprocessing carry far more of the load than in text-only training.

Real-world data, hundreds of terabytes

Demonstration data is collected on physical robots, not scraped from the web. Accumulated experiments quickly grow into hundreds of terabytes that have to move between robots and GPUs.

Simulation and RL, not just supervised passes

Much of robot learning is reinforcement learning in simulation: thousands of parallel environments, short bursty rollouts, then quiet while policies are validated on the robot. Models stay compact for on-device inference, so GPU demand arrives in waves, not a steady baseline.

Train in the cloud, act on the robot

Inference runs on the robot itself, where latency is a physical-safety problem. The cloud's job is training throughput, and fast checkpoint round-trips to get new models onto hardware.

Where GPU infrastructure gets in the way

The blockers robotics teams describe are rarely about raw GPU speed.

GPU access that can't keep up

Sweeps need eight GPUs today and two tomorrow. Waitlists, quota approvals, and allocations that fail on restart break the experiment cadence robot learning depends on.

Checkpoint and dataset round-trips

Every iteration uploads demonstration data and pulls checkpoints back down for on-robot testing. Storage and network quality, not GPU speed, often set the pace of iteration.

CPU bottlenecks nobody priced in

Image preprocessing and simulation rendering are CPU-hungry. A node with fast GPUs and starved CPUs still stalls the training loop.

Environment setup that breaks the run

Robotics stacks pin exact versions of CUDA, drivers, and tools like Isaac Sim, MuJoCo, ROS 2, and PhysX. One mismatch and the run doesn't slow down. It won't start. Teams can lose days to setup before a GPU does any useful work.

How VESSL Cloud fits

Compute plus workflow, mapped to what robot learning actually needs.

1

Self-serve GPUs, billed per second

Pick up GPUs when a sweep starts and release them when it ends, self-serve, with no quota tickets to file. When training becomes steady, reserved terms from three months guarantee capacity.

2

VM Cluster with root access

Scale to dedicated multi-node clusters: GPU nodes wired with InfiniBand, root SSH, and full driver control, as close to bare metal as a cloud gets.

3

Reproducible robotics environments

Pin your CUDA, simulator, and ROS 2 stack into a container image once. Every job and workspace starts from that exact environment, so a run that worked yesterday works today, and a teammate gets the same setup without rebuilding it.

4

Storage for robot-scale data

Object Storage holds hundreds-of-terabyte demonstration datasets at lower cost. Cluster Storage shares files across concurrent workloads with fast network performance. Attach volumes to any job.

5

One-command training sweeps

Define a job once with vesslctl and fan it out across the GPUs you need, with logs, metrics, and artifacts in a single view.

6

Workspaces for the dev loop

Connect VS Code or Cursor to a persistent GPU workspace, iterate interactively, and pause to keep the environment without paying for idle GPUs.

Physical AI on VESSL Cloud: FAQ

What GPUs do Physical AI teams train on?

Most robot foundation model training today runs on H100-class GPUs. On VESSL Cloud, H100 SXM starts at $3.19/hr self-serve, with H200 and Blackwell GPUs available on request. Inference usually runs on the robot's own edge hardware, so cloud spend stays focused on training.

How do I move hundreds of terabytes of robot data to the cloud?

Create an Object Storage volume, load data from the web console or the vesslctl CLI, and attach it to jobs at runtime. Cluster Storage covers fast file sharing between concurrent workloads.

Do I need a long-term contract to train on VESSL Cloud?

No. On-demand GPUs are billed per second with no commitment. When training becomes steady, reserved capacity starts at three-month terms with up to 15% off on-demand pricing.

Can I run Isaac Sim, MuJoCo, or ROS 2 on VESSL Cloud?

Yes. Simulators and middleware run as standard GPU jobs: pin your CUDA, simulator, and ROS 2 versions into a container image, and every job and workspace starts from that exact environment. Open VLA models such as NVIDIA's Isaac-GR00T fine-tune the same way, as multi-GPU jobs.

Where AI models
get their GPUs.

  • Start in minutes
  • Scale to multi-node clusters
  • Capacity reserved to your timeline
  • Dedicated support