Provisioned Throughput

Guaranteed throughput for open models

Contract the throughput you need in PTUs, and it holds even when the fleet is busy. VESSL Cloud runs the engines and the GPUs.

Holds at peak

contracted throughput, not best effort

Lower bill

below frontier API rates at sustained volume

One-line change

OpenAI- and Anthropic-compatible APIs

Monitoring

Usage and performance, right in VESSL Cloud

Compare peak usage with your contracted throughput, model by model. Click through the demo below.

What it solves

Throughput you can plan a product on

For coding agents, chat services, agent platforms, and internal AI workloads whose quality rides on tokens per minute.

Your contracted throughput holds at peak

Throughput, reliability, and time per output token (TPOT) targets are in your contract.

A lower cost per token than frontier APIs

When steady, high-volume traffic keeps your PTUs busy, they cost less per token than a comparable frontier model. The monthly bill follows your PTU count.

VESSL Cloud runs it for you

Keep your OpenAI or Anthropic client. Only the connection settings change.

The one line that changes

python
-client = OpenAI(api_key=KEY)
+client = OpenAI(api_key=KEY, base_url="https://inference.cloud.vessl.ai/v1")

Use your VESSL Cloud API key and the model ID from the model page.

Send your first request

PTU

A PTU is the unit you contract throughput in

A Provisioned Throughput Unit (PTU) is the same unit of serving capacity on every model.

1 PTU

Tokens it processes per minute, by model

Z.ai

GLM 5.3 Flash

tokens per minute

  • Input333,333
  • Cached input1,666,667
  • Output100,000

DeepSeek

DeepSeek V4.1 Flash

tokens per minute

  • Input166,667
  • Cached input7,142,857
  • Output41,667

One PTU covers any one of the three per minute. Real traffic mixes them, and a contract can set different values.

Need a different model? Tell us which model you need and we'll look into adding it.

How PTUs are billed

Pricing

One PTU rate for every model

Within your contracted capacity, you're billed for PTUs, not per token, and invoiced monthly.

$0.05

per PTU, per minute

$2,190

per PTU, per 730-hour month

Arranged with sales

PTU quantity and contract term

Monthly cost by workload: VESSL Cloud PTU (GLM 5.3 Flash) vs Claude Opus 5.5 and GPT-6.1 Sol

Illustrative scenario

We price the PTUs GLM 5.3 Flash needs at $0.05/min and compare them with the pay-as-you-go list rates of Claude Opus 5.5 and GPT-6.1 Sol, assuming continuous traffic for a 730-hour month. Both models charge these rates at every effort level, max included. The request profiles are in the assumptions.

VESSL Cloud PTUClaude Opus 5.5GPT-6.1 Sol

Coding Agent

  • VESSL Cloud PTU
    $11K
  • Claude Opus 5.5
    $155K14.1x
  • GPT-6.1 Sol
    $77K7.1x

Chat (single turn)

  • VESSL Cloud PTU
    $7K
  • Claude Opus 5.5
    $74K11.2x
  • GPT-6.1 Sol
    $37K5.6x

Agentic workflow (multi turn)

  • VESSL Cloud PTU
    $9K
  • Claude Opus 5.5
    $123K14.0x
  • GPT-6.1 Sol
    $61K7.0x

Tokens one GLM 5.3 Flash PTU processes per minute

InputOutputCached input
Tokens/min333,333100,0001,666,667

List price, $ per 1M tokens

InputOutputCached input
Claude Opus 5.5$4$20$0.20
GPT-6.1 Sol$2$10$0.10

Request profile per workload

WorkloadInput tok/reqOutput tok/reqCache hitCached tok/reqReq/sec
Coding Agent55,00030080%44,0001
Chat (single turn)1,00050020%2002
Agentic workflow (multi turn)20,0001,00070%14,0001

How each total adds up

  • Coding Agent: VESSL Cloud 5 PTU = $10,950Claude Opus 5.5 Input $115,632 + Cached input $23,126 + Output $15,768 = $154,526GPT-6.1 Sol Input $57,816 + Cached input $11,563 + Output $7,884 = $77,263
  • Chat (single turn): VESSL Cloud 3 PTU = $6,570Claude Opus 5.5 Input $21,024 + Cached input $0 + Output $52,560 = $73,584GPT-6.1 Sol Input $10,512 + Cached input $0 + Output $26,280 = $36,792
  • Agentic workflow (multi turn): VESSL Cloud 4 PTU = $8,760Claude Opus 5.5 Input $63,072 + Cached input $7,358 + Output $52,560 = $122,990GPT-6.1 Sol Input $31,536 + Cached input $3,679 + Output $26,280 = $61,495

The $0.05/min PTU rate is indicative and subject to change, cache hit rates and request profiles are assumptions, and the Claude Opus 5.5 and GPT-6.1 Sol rates are Anthropic's and OpenAI's standard public list prices, not Batch, Flex or Fast mode (OpenAI: prompts of 272K input tokens or fewer). Neither provider caches a prompt prefix under its minimum (512 tokens on Claude Opus 5.5, 1,024 on GPT-6.1 Sol), so a workload that reuses fewer tokens per request than that is billed uncached on that model. All three are costed on the same token counts. Higher effort levels spend more tokens at these same rates, and prompt caching adds write charges on both. Neither is counted here, so the Claude Opus 5.5 and GPT-6.1 Sol totals are a floor. Each token class is provisioned separately and rounded up to whole PTUs before the three are added. The purchase quantity is arranged individually with sales. This is a sizing illustration, not a quote.

Sources: Anthropic pricing, OpenAI pricing, Anthropic prompt caching, OpenAI prompt caching

Size your own traffic

Enter your peak traffic to see the PTUs it needs.

Start from

Model

80%
PTUs you need

5PTU

Estimated monthly cost

$10,950/mo

Tokens per minute at peak

Uncached input
660,000
Cached input
2,640,000
Output
18,000
Talk to salesSign in to VESSL Cloud

A planning figure, not a quote. Sized on the published per-PTU capacity at the $0.05/min rate over a 730-hour month; a contract can set a different capacity, and the PTU quantity and term are arranged with sales.

Performance

See how it runs on real requests

A real OpenCode agent session on the endpoint, with time to first token (TTFT) and time per output token (TPOT) measured on screen as the agent works.

The numbers on screen were measured in this session. They are not contracted targets.

Serving stack performance

vLLM recommended recipe = 1x

Our serving stack against the vLLM recommended recipe, on the same hardware and within the same latency limits.

  • 7.3x

    Peak tokens per minute

  • 2.4x

    Output tokens per second

  • 2.2x

    Input tokens per second

  • ~11x

    KV cache capacity

Kernel-level: public Triton kernel = 1x

2.22x

Prefill speed, our kernel vs the public Triton kernel

A larger KV cache keeps more context in memory, so first-token wait stays low as load rises and the gap widens under heavier traffic.

Both stacks ran on the same hardware within the same latency limits, under load modeled on real production traffic patterns. Every throughput figure is the peak each stack held while staying within those limits. The prefill figure is kernel-level, our kernel against the public Triton kernel on the same hardware.

Don't take our word for it

Ask for the harness we use to measure TTFT, TPOT, and token usage on our endpoint, with public datasets, published metric definitions, and raw per-request results, and run it on your own traffic.

Ask for the benchmark harness

How to start

Four steps to go live

We agree the terms first, then prove performance and cost on your real traffic in a proof of concept (PoC) before your endpoint goes live.

01

Consult

We profile your workload together: model, traffic, and TPOT target.

02

Contract

We fix the guaranteed throughput, the SLA, and the term.

03

Proof of concept

We tune for the agreed targets; you check performance and cost on your traffic.

04

Go live

Your production API key arrives at term start; set the base URL, key, and model ID.

What to bring to the consultation

  • Model

    The base model you want to serve

  • Traffic volume

    Expected daily tokens and peak tokens per minute (TPM)

  • Traffic shape

    Average input and output length per request, and expected cache hit rate

  • Target SLA

    Your required time per output token (TPOT)

Frequently Asked Questions

Can I use it today?

Yes. Provisioned Throughput is generally available. Talk to our team and we'll size your capacity, agree the terms, and run a proof of concept on your own traffic before you go live.

Which models can I serve?

We offer the latest models from open-model developers such as Z.ai and DeepSeek, and the lineup grows with demand, so ask about the model you need.

Which APIs does it support?

The endpoint supports both the OpenAI Chat Completions API and the Anthropic Messages API. Keep your existing client or SDK and change the connection settings. We hand those over when your endpoint is set up.

How is it priced?

$0.05 per PTU per minute, the same rate for every model, invoiced monthly. The PTU quantity and contract term are arranged individually with sales.

What does the SLA cover?

The SLA covers three things: the throughput your contracted PTUs guarantee, a minimum share of eligible requests completing successfully each month, and a time per output token (TPOT) target. The specific targets and the remedies for missing them are set per workload when you contract.

Do you store prompts and outputs?

By default, prompts and outputs are not kept beyond processing the request, and content logging runs only if you opt in. Limited exceptions for abuse, fraud, security, and legal requirements are set out in the service terms.

Is my data used to train models?

No. Inputs and outputs aren't used to train, fine-tune, or update any model without your written opt-in.

Which security certifications do you have?

VESSL AI is SOC 2 Type II certified. See our Trust Center at trust.vessl.ai for current controls and reports.

Where AI models
get their GPUs

  • Start in minutes
  • Scale to multi-node clusters
  • Capacity reserved to your timeline
  • Dedicated support