# VESSL AI > VESSL AI is the GPU Liquidity Layer — access, route, and guarantee NVIDIA GPU capacity across clouds globally. Rent A100, L40S, and H100 GPUs self-serve by the second with on-demand and reserved pricing; H200 and NVIDIA Blackwell (B200, GB200, B300) capacity is available on request. Dedicated multi-node H100 VM Clusters (beta) add InfiniBand, root SSH, and bare-metal-class performance for large-scale distributed training. VESSL AI provides unified GPU access across multiple cloud providers, so AI teams—from startups to large enterprises—can run training, fine-tuning, and inference without waitlists, quota limits, or vendor lock-in. The company positions itself as "the orchestration layer for AI infrastructure," unifying fragmented compute into a single interface. VESSL Cloud launched in February 2026. ## Quick Facts - **GPU Models**: NVIDIA A100, L40S, H100 (self-serve on-demand); H200, B200, GB200, B300 (Blackwell Ultra) on request - **Pricing Tiers**: On-Demand (per-second, self-serve), Reserved (guaranteed capacity, up to 15% off, terms from 3 months), Spot (preemptible with auto-checkpointing — rolling out) - **VM Clusters (beta)**: dedicated H100 SXM clusters, 1–8 nodes self-serve, with InfiniBand and root SSH - **Certifications**: SOC 2 Type II - **Key Features**: Multi-cloud failover, auto-checkpointing, real-time monitoring, VS Code integration, multi-cluster management, InfiniBand multi-node training - **Founded**: 2020 · **Customers**: 100+ · **Offices**: San Francisco & Seoul ## What VESSL AI Does VESSL AI is a GPU cloud that aggregates capacity across multiple providers and routes workloads to wherever GPUs are available. Instead of managing quotas, procurement, and per-cloud lock-in, teams pick a GPU and start in minutes. - **Run your way** — choose spot, on-demand, or reserved capacity across A100, H100, H200, B200, GB200, and B300. Mix and match to fit workload and budget. - **Scale from 1 to hundreds of GPUs** — scale to zero when idle, then back up to hundreds when needed, without losing your environment. - **Multi-cloud failover** — workloads reroute automatically across providers to keep training and inference running. - **IDE-native** — connect VS Code directly to remote GPUs; no context switching. - **Multi-cluster** — unified management and a single view across regions and providers. ## Main Pages - [Home](https://vessl.ai): Platform overview and value proposition - [Pricing](https://vessl.ai/en/pricing): GPU pricing — On-Demand, Reserved, and Spot - [H100 GPU Cloud](https://vessl.ai/en/gpu/h100): NVIDIA H100 SXM (80GB HBM3) on-demand from $2.39/hr — self-serve, no waitlist - [H200 GPU Cloud](https://vessl.ai/en/gpu/h200): NVIDIA H200 SXM (141GB HBM3e) for long-context and large-model inference — talk to sales - [B200 GPU Cloud](https://vessl.ai/en/gpu/b200): NVIDIA Blackwell B200 (192GB HBM3e, FP4) reserved capacity — talk to sales - [B300 GPU Cloud](https://vessl.ai/en/gpu/b300): NVIDIA B300 (Blackwell Ultra, up to 288GB HBM3e) reserved capacity — talk to sales - [H100 vs H200 comparison](https://vessl.ai/en/gpu/h100-h200): NVIDIA Hopper (H100, H200) side-by-side, with on-page pricing - [B200 vs B300 comparison](https://vessl.ai/en/gpu/b200-b300): NVIDIA Blackwell (B200, B300) side-by-side — talk to sales - [Talk to Sales](https://vessl.ai/talk-to-sales): Reserved capacity, enterprise, and Blackwell inquiries - [Careers](https://vessl.ai/en/careers): Open roles and company overview - [English Site](https://vessl.ai/en) · [Korean Site](https://vessl.ai/ko) ## Products ### GPU Cloud - **On-Demand**: Self-serve, pay-per-second GPUs with automatic multi-cloud failover. Live self-serve pricing for A100 SXM, L40S, and H100 SXM. Best for production workloads and interactive development. - **Reserved**: Guaranteed capacity with up to 15% off on-demand pricing; terms from 3 months; dedicated support and volume discounts. Best for large-scale training and mission-critical AI. Available for A100, H100, L40S, B200, GB200, and B300. - **Spot**: Preemptible capacity at steep discounts (up to 90% off) with auto-checkpointing for interrupted runs. Best for research, batch jobs, and experimentation. (Rolling out.) ### VM Clusters (beta) Dedicated, single-tenant GPU clusters for multi-node distributed training: - Rent **1–8 nodes** of **8× NVIDIA H100 SXM** each, self-serve in VESSL Cloud; 9+ nodes or other GPUs (H200, B200, GB200, B300) through sales. - Delivered as **VMs with GPU passthrough** for bare-metal-class performance — single-tenant, with no other customers sharing your hardware. - **Root SSH** on every node, **InfiniBand** between nodes by default, a dedicated public IP per node, a 2 TiB boot disk, and optional NFS shared storage. - **Bring your own stack** — install Slurm, Kubernetes, security agents, or custom drivers on a standard Ubuntu host with full root access. - Billed as **prepaid reserved contracts** (H100 from $2.39/GPU/hr on-demand, with reserved discounts). Most clusters are ready within a few business days of ordering. ### Storage - **Cluster Storage**: High-performance shared storage across workloads, with fast network performance for collaborative training. - **Object Storage**: Cost-effective storage for large datasets and artifacts. - **NFS shared storage** (VM Clusters): optional shared filesystem mounted across all nodes in a cluster. ## GPU Pricing & Availability On-demand, per-GPU pricing as listed on the [pricing page](https://vessl.ai/en/pricing): | GPU | VRAM | On-Demand (per GPU) | Access | |-----|------|---------------------|--------| | NVIDIA A100 SXM | 80GB | $1.55/hr | Self-serve | | NVIDIA L40S | 48GB | $1.80/hr | Self-serve | | NVIDIA H100 SXM | 80GB | $2.39/hr | Self-serve | | NVIDIA H200 SXM | 141GB | On request | Talk to sales | | NVIDIA B200 | 192GB | On request | Talk to sales | | NVIDIA GB200 | 192GB | On request | Reserved / sales | | NVIDIA B300 (Blackwell Ultra) | up to 288GB | On request | Talk to sales | Reserved commitments save up to 15% vs on-demand (terms from 3 months). Academic discounts are available for universities and research labs. Prices are indicative and may change—see the pricing page for current rates. ## Technical Specifications (Hopper & Blackwell) Verified against NVIDIA primary sources (2026-05): | Spec | H100 SXM | H200 SXM | B200 | B300 (Blackwell Ultra) | |------|----------|----------|------|------------------------| | Architecture | Hopper | Hopper | Blackwell | Blackwell | | Memory | 80GB HBM3 | 141GB HBM3e | 192GB HBM3e | up to 288GB HBM3e | | Memory bandwidth | 3.35 TB/s | 4.8 TB/s | 8 TB/s | 8 TB/s | | NVLink | 900 GB/s | 900 GB/s | 1.8 TB/s | 1.8 TB/s | | FP8 (peak, sparse) | 3,958 TFLOPS | 3,958 TFLOPS | 9 PFLOPS | 10 PFLOPS | | FP4 (peak, sparse) | — | — | 18 PFLOPS | 20 PFLOPS | | Max TDP/TGP | 700W | 700W | up to 1,000W | up to 1,400W | | GPUs per node | 8 (HGX H100) | 8 (HGX H200) | 8 (HGX B200) | 8 (HGX B300) | *FP throughput figures are peak performance with sparsity (dense is roughly half), per NVIDIA datasheets. Final specs may vary by node configuration. A100 (80GB) and L40S (48GB) are also available on-demand; GB200 pairs a Grace CPU with Blackwell GPUs at 192GB per GPU. ## Interfaces - **Web Console**: Visual cluster management at cloud.vessl.ai - **CLI**: `vessl run` native workflows - **VS Code integration**: Connect your IDE directly to remote GPUs - **Auto Failover**: Seamless provider switching - **Multi-Cluster**: Unified view across regions and providers ## Use Cases - LLM post-training and fine-tuning - Large-scale inference services - Physical AI — robotics and autonomous driving - AI for Science and academic research - Multi-node distributed training (InfiniBand) on reserved or VM Cluster capacity ## Target Customers ### AI Startups No quota limits or waitlists. Multi-cloud failover built-in. Pay-as-you-go pricing. Scale from prototype to production across H100, A100, H200, B200, GB200, and B300. ### Enterprise AI Teams Guaranteed capacity with enterprise-grade reliability. SOC 2 Type II. Dedicated account management and onboarding. Custom integrations and on-premise support. Volume discounts at scale. ### Research & Academia Self-serve GPU access in minutes. VS Code and CLI-native workflows. Spot instances for budget-conscious labs. Auto-checkpointing for preemptible runs. Discounted academic rates. ## Customer Success Stories - **UC Berkeley (BAIR)**: Shifted time from infrastructure wrangling to experiment design, with fire-and-forget reliability. - **NYU**: Eliminated GPU queue waits from shared Slurm clusters; same-day experiment iteration and improved reproducibility across collaborators. - **Hanwha Life**: Infrastructure setup time cut ~90% (months to instant); 3x faster project kickoff. - **Tmap Mobility**: Multi-cloud routing cut compute costs and improved training reliability across regions. - **Hyundai Motor Company**: Integrated 100TB+ of multi-region data and cut model deployment time from five months to one week. - **NXN AI**: GPU infrastructure TCO down 50%; monthly experiments up 200%. - **Tomorro Robotics**: Standardized a ROS 2 stack on VESSL Workspace and Storage, boosting collaboration ~80%. Also trusted by KT, GS Retail, Yanolja, Upstage, Rebellions, and Scatter Lab, plus research groups at Stanford, MIT, CMU, Seoul National University, KAIST, and the University of Washington. ## Resources - [Documentation](https://docs.cloud.vessl.ai): Guides and API reference - [Blog](https://vessl.ai/en/blog): Engineering stories and best practices - [Changelog](https://docs.cloud.vessl.ai/changelog/overview/updates): Latest feature releases - [Trust Center](https://trust.vessl.ai): Security controls and SOC 2 Type II compliance - [Brand Assets](https://vesslai.notion.site/VESSL-AI-Brand-Assets-a07ff2fc35344bd1bfb17aa7d88d87e5): Press kit and logos - [Events](https://lu.ma/user/vesslai): Talks, workshops, and meetups ## Connect - GitHub: https://github.com/vessl-ai/ - LinkedIn: https://www.linkedin.com/company/vesslai/ - X: https://x.com/vesslai - YouTube: https://www.youtube.com/channel/UCPCXK0wXfBWSEmMW2yEbt5w ## Company - **Entity**: VESSL AI, Inc. - **Founded**: 2020 - **Offices**: San Francisco, CA, USA & Seoul, South Korea - **Customers**: 100+ across enterprise, startups, government, and academia - **Total funding**: $16M+ - **Milestone**: VESSL Cloud launched February 2026 - **Contact**: [Talk to Sales](https://vessl.ai/talk-to-sales) ## FAQ **Q: Which GPUs can I start using self-serve right now?** A: A100 SXM (80GB, $1.55/hr), L40S (48GB, $1.80/hr), and H100 SXM (80GB, $2.39/hr) are available on-demand and self-serve at cloud.vessl.ai—no waitlist. H200 and Blackwell (B200, GB200, B300) capacity is available on request. **Q: How do I get NVIDIA Blackwell (B200/B300) GPUs?** A: Blackwell capacity is allocated on request. Talk to sales and we'll secure B200 or B300 (Blackwell Ultra) capacity matched to your timeline. We provision HGX B200/B300 nodes (8 GPUs each) with high-speed InfiniBand, from a single node to large multi-node clusters. **Q: What are VM Clusters?** A: VM Clusters (beta) let you rent a dedicated cluster of H100 nodes (1–8 nodes self-serve, 8× H100 SXM each) as VMs with GPU passthrough—root SSH, InfiniBand by default, and full kernel/driver control for multi-node distributed training. Larger clusters (9+ nodes) or other GPUs are available through sales. **Q: What's the difference between Spot, On-Demand, and Reserved?** A: On-Demand is reliable, pay-per-second capacity with automatic failover. Reserved guarantees capacity with up to 15% off (terms from 3 months), dedicated support, and volume discounts. Spot uses preemptible excess capacity at steep discounts with auto-checkpointing. **Q: What's the difference between the H100 and H200?** A: Both share the same Hopper compute, but the H200 carries 141GB of faster HBM3e memory (vs 80GB HBM3 on the H100) at 4.8 TB/s bandwidth—fitting larger models, bigger batches, and longer context windows. **Q: Do you offer academic pricing?** A: Yes. Universities and research labs can access discounted rates and flexible terms—contact us for details. **Q: How does multi-cloud failover work?** A: If your primary cloud provider experiences issues, VESSL automatically reroutes your workloads to available capacity on another provider—no manual intervention required.