# VESSL AI > VESSL AI is the AI Cloud for All Models: access, route, and guarantee NVIDIA GPU capacity across clouds globally. Rent A100, L40S, and H100 GPUs self-serve by the second with on-demand and reserved pricing; H200 and NVIDIA Blackwell (B200, B300, GB300) run as reserved VM Cluster capacity. Dedicated multi-node H100 VM Clusters (beta) add InfiniBand, root SSH, and bare-metal-class performance for large-scale distributed training. VESSL AI provides unified GPU access across multiple cloud providers, so AI teams, from startups to large enterprises, can run training, fine-tuning, and inference without quota limits or vendor lock-in. The company's tagline is "Where AI models get their GPUs." The platform unifies fragmented compute into a single interface. VESSL Cloud launched in February 2026. ## Quick Facts - **GPU Models**: NVIDIA A100, L40S, H100 (self-serve on-demand); H200, B200, B300, GB300 (reserved via VM Cluster) - **Pricing Tiers**: On-Demand (per-second, self-serve) and Reserved (guaranteed capacity, up to 15% off, terms from 3 months) - **VM Clusters (beta)**: dedicated H100 SXM clusters, 1–8 nodes self-serve, with InfiniBand and root SSH - **Certifications**: SOC 2 Type II - **Key Features**: Multi-cloud GPU access, multi-cluster management, real-time monitoring, VS Code integration, InfiniBand multi-node training - **Founded**: 2020 · **Customers**: 100+ · **Offices**: San Francisco & Seoul ## What VESSL AI Does VESSL AI is a GPU cloud that aggregates capacity across multiple providers and routes workloads to wherever GPUs are available. Instead of managing quotas, procurement, and per-cloud lock-in, teams pick a GPU and start in minutes. - **Run your way**: A100, L40S, and H100 on-demand by the second; H200, B200, B300, and GB300 as reserved VM Cluster capacity. Mix and match to fit workload and budget. - **Scale from 1 to hundreds of GPUs**: scale to zero when idle, then back up to hundreds when needed, without losing your environment. - **Multi-cloud access**: source and manage GPU capacity across multiple providers from one interface. - **IDE-native**: connect VS Code directly to remote GPUs; no context switching. - **Multi-cluster**: unified management and a single view across regions and providers. ## Main Pages - [Home](https://vessl.ai): Platform overview and value proposition - [VESSL Cloud](https://vessl.ai/en/cloud): Product tour covering GPU workspaces, one-command jobs (`vesslctl`), multi-node VM Clusters, and storage in one platform; run training, fine-tuning, and inference workloads - [Storage](https://vessl.ai/en/storage): Two storage tiers for GPU workloads: a warm tier (Cluster Storage, $0.20/GiB-month, CephFS on NVMe, cluster-local) and a cold tier (Object Storage, $0.03/GiB-month, S3-backed, cross-cluster), billed daily on actual usage; create and manage volumes from the web console or `vesslctl volume` - [VM Cluster](https://vessl.ai/en/vm-clusters): Dedicated multi-node GPU clusters wired with InfiniBand, from NVIDIA H100 through Blackwell; 8 GPUs per node, up to 8 nodes (64 GPUs), full root access on every node; talk to sales - [Inference Endpoints](https://vessl.ai/en/inference): Dedicated inference endpoints with Anthropic- and OpenAI-compatible APIs and transparent token billing; coming soon, design partners welcome - [Use cases](https://vessl.ai/en/use-cases): How AI teams put VESSL Cloud to work, with deep dives per workload - [Physical AI use case](https://vessl.ai/en/use-cases/physical-ai): GPUs for robot foundation models and RL in simulation - [Foundation models & NLP use case](https://vessl.ai/en/use-cases/nlp): Full-stack multi-node GPU clusters for large-model training - [Academia use case](https://vessl.ai/en/use-cases/academia): Self-serve GPUs and reproducible environments for AI research labs - [Pricing](https://vessl.ai/en/pricing): GPU pricing (On-Demand and Reserved) - [H100 GPU Cloud](https://vessl.ai/en/gpu/h100): NVIDIA H100 SXM (80GB HBM3) on-demand from $2.98/hr, self-serve - [H200 GPU Cloud](https://vessl.ai/en/gpu/h200): NVIDIA H200 SXM (141GB HBM3e) for long-context and large-model inference; talk to sales - [B200 GPU Cloud](https://vessl.ai/en/gpu/b200): NVIDIA Blackwell B200 (192GB HBM3e, FP4) reserved capacity; talk to sales - [B300 GPU Cloud](https://vessl.ai/en/gpu/b300): NVIDIA B300 (Blackwell Ultra, up to 288GB HBM3e) reserved capacity; talk to sales - [H100 vs H200 comparison](https://vessl.ai/en/gpu/h100-h200): NVIDIA Hopper (H100, H200) side-by-side, with on-page pricing - [B200 vs B300 comparison](https://vessl.ai/en/gpu/b200-b300): NVIDIA Blackwell (B200, B300) side-by-side; talk to sales - [Talk to Sales](https://vessl.ai/talk-to-sales): Reserved capacity, enterprise, and Blackwell inquiries - [Careers](https://vessl.ai/en/careers): Open roles and company overview - [English Site](https://vessl.ai/en) · [Korean Site](https://vessl.ai/ko) ## Products ### GPU Cloud - **On-Demand**: Self-serve, pay-per-second GPUs sourced across multiple clouds. Live self-serve pricing for A100 SXM, L40S, and H100 SXM. Best for production workloads and interactive development. - **Reserved**: Guaranteed capacity with up to 15% off on-demand pricing; terms from 3 months; dedicated support and volume discounts. Best for large-scale training and mission-critical AI. Available for A100, H100, and L40S; H200, B200, B300, and GB300 run exclusively as reserved VM Cluster capacity. ### VM Clusters (beta) Dedicated, single-tenant GPU clusters for multi-node distributed training: - Rent **1–8 nodes** of **8× NVIDIA H100 SXM** each, self-serve in VESSL Cloud; 9+ nodes or other GPUs (H200, B200, GB200, B300) through sales. - Delivered as **VMs with GPU passthrough** for bare-metal-class performance: single-tenant, with no other customers sharing your hardware. - **Root SSH** on every node, **InfiniBand** between nodes by default, a dedicated public IP per node, a 2 TiB boot disk, and optional NFS shared storage. - **Bring your own stack**: install Slurm, Kubernetes, security agents, or custom drivers on a standard Ubuntu host with full root access. - Billed as **prepaid reserved contracts** with volume and term discounts. Most clusters are ready within a few business days of ordering. ### Storage - **Cluster Storage** (warm tier): High-availability CephFS-on-NVMe shared storage across workloads in a cluster, with fast cluster-local network performance for collaborative training. Read-write-many; data survives workspace stop and terminate. - **Object Storage** (cold tier): Cost-effective S3-backed storage for datasets, checkpoints, logs, and long-term artifacts, accessible from any cluster in the organization. - **NFS shared storage** (VM Clusters): optional shared filesystem mounted across all nodes in a cluster. ## GPU Pricing & Availability Per-GPU hourly rates as listed on the [pricing page](https://vessl.ai/en/pricing): | GPU | VRAM | Hourly rate (per GPU) | Access | |-----|------|---------------------|--------| | NVIDIA A100 SXM | 80GB | $1.48/hr | Self-serve | | NVIDIA L40S | 48GB | $1.80/hr | Self-serve | | NVIDIA H100 SXM | 80GB | $2.98/hr | Self-serve | | NVIDIA H200 SXM | 141GB | $4.38/hr | Talk to sales | | NVIDIA B200 | 192GB | $5.88/hr | Talk to sales | | NVIDIA B300 (Blackwell Ultra) | 288GB | $7.38/hr | Talk to sales | | NVIDIA GB300 | 288GB per GPU | Reserved only | Talk to sales | Reserved commitments on self-serve GPUs (A100, L40S, H100) save up to 15% (terms from 3 months); H200 and Blackwell reserved contracts are quoted by term and volume with sales. Prices are indicative and may change. See the pricing page for current rates. ## Technical Specifications (Hopper & Blackwell) Verified against NVIDIA primary sources (2026-05): | Spec | H100 SXM | H200 SXM | B200 | B300 (Blackwell Ultra) | |------|----------|----------|------|------------------------| | Architecture | Hopper | Hopper | Blackwell | Blackwell | | Memory | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | | Memory bandwidth | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | | NVLink | 900 GB/s | 900 GB/s | 1.8 TB/s | 1.8 TB/s | | FP8 (peak, sparse) | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | | FP4 (peak, sparse) | N/A | N/A | 18 PFLOPS | 20 PFLOPS | | Max TDP/TGP | 700W | 700W | up to 1,000W | up to 1,400W | | GPUs per node | 8 (HGX H100) | 8 (HGX H200) | 8 (HGX B200) | 8 (HGX B300) | *FP throughput figures are peak performance with sparsity (dense is roughly half), per NVIDIA datasheets. Final specs may vary by node configuration. A100 (80GB) and L40S (48GB) are also available on-demand; GB300 pairs a Grace CPU with Blackwell Ultra GPUs at 288GB per GPU. ## Interfaces - **Web Console**: Visual cluster management at cloud.vessl.ai - **CLI**: `vesslctl` native workflows - **VS Code integration**: Connect your IDE directly to remote GPUs - **Multi-cloud access**: GPU capacity sourced across multiple providers - **Multi-Cluster**: Unified view across regions and providers ## Use Cases - LLM post-training and fine-tuning - Large-scale inference services - Physical AI: robotics and autonomous driving - AI for Science and academic research - Multi-node distributed training (InfiniBand) on reserved or VM Cluster capacity ## Target Customers ### AI Startups No quota limits. Multi-cloud GPU access built-in. Pay-as-you-go pricing. Scale from prototype to production, from on-demand A100 and H100 to reserved Blackwell capacity. ### Enterprise AI Teams Guaranteed capacity with enterprise-grade reliability. SOC 2 Type II. Dedicated account management and onboarding. Custom integrations and dedicated cluster options. Volume discounts at scale. ### Research & Academia Self-serve GPU access in minutes. VS Code and CLI-native workflows. Per-second, pay-as-you-go billing with no idle spend. ## Customer Success Stories - **UC Berkeley (BAIR)**: Shifted time from infrastructure wrangling to experiment design, with fire-and-forget reliability. - **NYU**: Eliminated GPU queue waits from shared Slurm clusters; same-day experiment iteration and improved reproducibility across collaborators. - **Hanwha Life**: Infrastructure setup time cut ~90% (months to instant); 3x faster project kickoff. - **Tmap Mobility**: Multi-cloud routing cut compute costs and improved training reliability across regions. - **Hyundai Motor Company**: Integrated 100TB+ of multi-region data and cut model deployment time from five months to one week. - **NXN AI**: GPU infrastructure TCO down 50%; monthly experiments up 200%. - **Tomorro Robotics**: Standardized a ROS 2 stack on VESSL Workspace and Storage, boosting collaboration ~80%. Also trusted by KT, GS Retail, Yanolja, Upstage, Rebellions, and Scatter Lab, plus research groups at Stanford, MIT, CMU, Seoul National University, KAIST, and the University of Washington. ## Resources - [Documentation](https://docs.cloud.vessl.ai): Guides and API reference - [Blog](https://vessl.ai/en/blog): Engineering stories and best practices - [Changelog](https://docs.cloud.vessl.ai/changelog/overview/updates): Latest feature releases - [Trust Center](https://trust.vessl.ai): Security controls and SOC 2 Type II compliance - [Legal](https://docs.cloud.vessl.ai/legal/overview): Terms of service, privacy policy, and other legal documents - [Brand Assets](https://vesslai.notion.site/VESSL-AI-Brand-Assets-a07ff2fc35344bd1bfb17aa7d88d87e5): Press kit and logos - [Events](https://lu.ma/user/vesslai): Talks, workshops, and meetups ## Connect - GitHub: https://github.com/vessl-ai/ - LinkedIn: https://www.linkedin.com/company/vesslai/ - X: https://x.com/vesslai - YouTube: https://www.youtube.com/channel/UCPCXK0wXfBWSEmMW2yEbt5w ## Company - **Entity**: VESSL AI, Inc. - **Founded**: 2020 - **Offices**: San Francisco, CA, USA & Seoul, South Korea - **Customers**: 100+ across enterprise, startups, government, and academia - **Total funding**: $16M+ - **Milestone**: VESSL Cloud launched February 2026 - **Contact**: [Talk to Sales](https://vessl.ai/talk-to-sales) ## FAQ **Q: Which GPUs can I start using self-serve right now?** A: A100 SXM (80GB, $1.48/hr), L40S (48GB, $1.80/hr), and H100 SXM (80GB, $2.98/hr) are available on-demand and self-serve at cloud.vessl.ai. H200 and Blackwell (B200, B300, GB300) run as reserved VM Cluster capacity. Talk to sales to reserve. **Q: How do I get NVIDIA Blackwell (B200/B300) GPUs?** A: Blackwell capacity is allocated on request. Talk to sales and we'll secure B200 or B300 (Blackwell Ultra) capacity matched to your timeline. We provision HGX B200/B300 nodes (8 GPUs each) with high-speed InfiniBand, from a single node to large multi-node clusters. **Q: What are VM Clusters?** A: VM Clusters (beta) let you rent a dedicated cluster of H100 nodes (1–8 nodes self-serve, 8× H100 SXM each) as VMs with GPU passthrough: root SSH, InfiniBand by default, and full kernel/driver control for multi-node distributed training. Larger clusters (9+ nodes) or other GPUs are available through sales. **Q: What's the difference between On-Demand and Reserved?** A: On-Demand is reliable, pay-per-second capacity sourced across multiple clouds, self-serve for A100, L40S, and H100. Reserved guarantees capacity with up to 15% off (terms from 3 months), dedicated support, and volume discounts. H200 and Blackwell GPUs run only as reserved VM Cluster capacity. **Q: What's the difference between the H100 and H200?** A: Both share the same Hopper compute, but the H200 carries 141GB of faster HBM3e memory (vs 80GB HBM3 on the H100) at 4.8 TB/s bandwidth, fitting larger models, bigger batches, and longer context windows. **Q: How does multi-cloud access work?** A: VESSL sources GPU capacity across multiple cloud providers and routes workloads to wherever GPUs are available, all managed from one interface. You pick a GPU and start in minutes instead of negotiating quota with a single provider.