How LLM providers sell reserved capacity

Nine ways to buy inference capacity before you use it, put in one set of columns, with every term read from the vendor's own page.

By Standard Thinking · Updated

Nine offerings sell, or until recently sold, inference capacity before it is used, and five publish a price you can read without a sales call. They agree on the idea and differ on the rest: what a unit measures, the meter, the term, and what happens above the reservation. Every cell was read from the vendor's own page on September 16, 2026.

The nine offerings in one set of columns

Offering Unit Meter Minimum term Price public Above it
Azure AI Foundry provisioned throughput PTU Hourly None Partial Spillover
Amazon Bedrock Provisioned Throughput Model Unit Hourly None No Not stated
Vertex AI Provisioned Throughput GSU Hourly 1 week Yes PAYG
OpenAI Scale Tier Token unit Per day 30 days Yes PAYG
Anthropic Priority Tier, retired Capacity commitment Not stated 1 month No Fallback
Together AI PTUs and hardware PTU or GPU Per minute Not stated Partial Not stated
Fireworks AI on-demand deployments GPU Per second None Yes Hardware cap
Baseten dedicated deployments GPU instance Per minute None Yes Hardware cap
Product Plan Output tokens per second Monthly 1 month Yes Extra + blocks

Notes by provider

  • Azure. Deployed PTUs are billed "regardless of the number of tokens consumed." Throughput per PTU varies by model, each model has a minimum deployment size, spillover excludes Azure DeepSeek and Meta Llama, and the page read names no rate.
  • Bedrock. One Model Unit specifies input and output tokens per minute across all requests, though the specification and the price come from an AWS account manager. A one-month or six-month commitment cannot be deleted before it ends, and billing continues until you delete it.
  • Vertex AI. Published for global endpoints at $7.14 per GSU-hour on a one-week commit, down to $2.739726027 on one year. Overages are controllable per request, and cache discounts arrive as a reduced burndown rate rather than a separate price.
  • OpenAI. Published per model snapshot, per unit per day: GPT-5.5 at 50,000 TPM for $750.00. Billing TPM is averaged over 15-minute intervals aligned to the hour, and excess runs at pay-as-you-go rates on Fast mode, automatic since July 2026.
  • Anthropic. Closed: "Priority Tier capacity commitments are no longer available for purchase." A commitment set input and output tokens per minute for one model version, with cache reads counted at 0.1 tokens.
  • Together AI. Hardware is published at $5.49 per HGX H100 GPU-hour and $8.99 per HGX B200. The PTU table lists input, cached, and output TPM per PTU, but its per-model rows did not render, so no rate is quoted.
  • Fireworks AI. Start-up time is not charged, and no token entitlement is published. Prices run $8.00 an hour for an H100 80 GB to $20.00 for a GB300 288 GB, in the column headed "from Sep 1."
  • Baseten. Idle time is not billed and scaling is the customer's to configure: $0.10833 a minute for an H100 80 GiB, $0.16633 for a B200, on an entry plan of $0 a month, pay as you go.
  • Product Plan. $1,200 per unit per model per month, $3,000 for one allocation of all three, and $5 per extra unit per two-hour period for authorized auto scaling, subject to availability. Capacity above the base is not reserved: free while the shared pool has room.

What a unit measures differs more than the price

Four kinds of unit are in that table; none converts into another. Tokens per minute is the common shape: Bedrock and Anthropic size input and output separately, Azure and Together sell one unit whose throughput depends on the model, and OpenAI sells separate bundles before GPT-5.4 and a combined one after. Google counts input plus output per second: burndown rates put Gemini 3.8 Flash at 675 tokens per second per GSU, gpt-oss 120B at 11,205. Product Plan counts output tokens per second: 40 on Kimi K3, 100 on GLM-5.3, 300 on DeepSeek V4.1 Flash, with input, concurrency, burst length, and latency set in the order, and no monthly token cap. The fourth unit is a GPU, promising nothing about tokens.

The meter is not the commitment

A month-shaped price is unusual: Google publishes a monthly view of an hourly rate, and Product Plan bills monthly with peaks in two-hour periods. An Azure reservation discounts the PTU billing meter for one month or one year, and "Reservations don't guarantee capacity" — create the deployment first, then buy the reservation. Google binds harder: "you can't cancel the order in the middle of your term," and weekly terms cannot renew automatically.

Who is allowed to buy

OpenAI's Scale Tier "is available to Enterprise customers" after an order form is signed, Bedrock sends the unit specification and the price to an account manager, Together quotes reserved hardware through sales, and Anthropic's door is shut. Google, Fireworks, Baseten, and Product Plan publish a number you can read first, though Google processes GSU orders against available capacity, which "might take from a few minutes to a few weeks." Seven of the nine serve open models; Bedrock lists Llama 3.1 and 3.2, Google's in Preview.

What to ask before you reserve

  1. Restate every unit as answers per hour at your own answer length before comparing prices; the capacity calculator does that arithmetic for the monthly shape.
  2. What happens above the reservation? Three behaviors appear: a pay-as-you-go fallback, a hardware ceiling with nothing beyond it, and capacity that is not reserved.
  3. How long are you committed? Minimum terms run from none to a month and optional commitments reach a year; read the exit clause next to the discount, not instead of it.
  4. Is the published price current? Google's 50% Flash credit ends December 31, 2026, Fireworks' table carries two dated columns, and Together has run dated promotions on its H100 rate.

This page restates published terms in one set of columns and recommends no offering. Product Plan is available now: log in to order a base, or see it costed against pay per token in buying inference on a monthly plan.

Sources and verification

Sources used for this resource. Verification dates describe our source checks, not provider effective dates.

Azure AI Foundry provisioned throughput documentation

Provisioned deployments are "billed at an hourly rate ($/PTU/hr) based on the number of PTUs deployed, regardless of the number of tokens consumed"; Azure Reservations exchange a 1-month or 1-year commitment for a discounted effective rate, and "Reservations don't guarantee capacity"; the tokens per minute a given PTU count delivers varies by model and each model carries a minimum deployment size; spillover routes overflow requests to a standard deployment for Azure OpenAI models, while "Foundry Models from other providers (Azure DeepSeek, Meta Llama) don't currently support spillover". No $/PTU/hr figure appears on this page.

Checked

Amazon Bedrock Provisioned Throughput user guide

"You're billed hourly for a Provisioned Throughput that you purchase"; a Model Unit specifies the input tokens it can process and the output tokens it can generate across all requests within one minute; commitment levels are none, 1 month, or 6 months, with no deletion before the term ends and "billing continues until you delete the Provisioned Throughput"; the MU specification and the price per MU are directed to an AWS account manager.

Checked

Amazon Bedrock, supported Regions and models for Provisioned Throughput

The model list available for Provisioned Throughput purchases, which includes the open-weight Meta Llama 3.1 70B and 8B Instruct and Llama 3.2 1B, 3B, 11B, and 90B Instruct entries, each shown with us-west-2 single-region support.

Checked

Vertex AI Provisioned Throughput overview

Provisioned Throughput described as "a fixed-cost, fixed-term subscription available in several term-lengths that reserves throughput for supported generative AI models", bought for deterministic costs "by paying a fixed monthly or weekly price with control of overages".

Checked

Vertex AI, calculate Provisioned Throughput requirements

"A Generative AI Scale Unit (GSU) is a measure of throughput for your prompts and responses"; a burndown rate "converts the input and output units ... to input tokens per second" to produce a standard unit across models; unused throughput does not carry over to the next month; the cache discount is applied as a reduced burndown rate, so 1,000 cached tokens burn down 100 tokens per second on the page's Gemini example.

Checked

Vertex AI, models supported by Provisioned Throughput

"Your per-second throughput is defined as your prompt input and generated output across all requests per second"; Gemini 3.8, 3.7, and 3.6 Flash are listed at 675 tokens per second per GSU with an output text token burning 5 tokens; a separate table covers "open models that support Provisioned Throughput" in Preview, including OpenAI gpt-oss 120B at 11,205 tokens per second per GSU, GLM 5, Qwen3-Next-80B, Llama, Kimi K2 Thinking, MiniMax M2, and DeepSeek-V3.2.

Checked

Vertex AI, purchase Provisioned Throughput

"you can't cancel the order in the middle of your term"; monthly subscriptions can renew automatically while "Weekly terms don't support automatic renewal"; "If your throughput exceeds your Provisioned Throughput order amount, overages are processed and billed as standard pay-as-you-go. You can control overages on a per-request basis"; orders are processed against available capacity and "might take from a few minutes to a few weeks".

Checked

Google Cloud Agent Platform pricing, Provisioned Throughput section

Published price per GSU for global endpoints: $7.14 per hour on a 1-week commit, $3.698630137 on 1 month, $3.287671233 on 3 months, and $2.739726027 on 1 year, with higher non-global rates; the page's own worked example converts these to $1,200 per GSU per week and $2,700, $2,400, and $2,000 per GSU per month. A 50% monthly billing credit on Provisioned Throughput spending for three Gemini Flash models runs from August 13 to December 31, 2026.

Checked

OpenAI Scale Tier for API customers

"Scale Tier lets you purchase a set number of API input and output tokens per minute (known as "token units") upfront for access to one specific model snapshot. Each token unit is purchased for a minimum of 30 days."; published per-model prices per unit per day, such as GPT-5.5 at 50,000 TPM for $750.00 and GPT-4.1 at 30,000 input TPM for $110.00 with a 2,500 output TPM bundle at $36.00; combined bundles on GPT-5.4 and GPT-5.5 where "one output token counts as 6"; billing TPM averaged over 15-minute intervals aligned to the top of the hour, with excess "billed at pay-as-you-go (PAYG) rates on Fast mode"; "This offering is available to Enterprise customers." A successor, Reserved Tier, is named for GPT-5.6 and later models.

Checked

OpenAI Help Center, Scale Tier for existing enterprise customers

Confirms the same token-unit definition and 30-day minimum, and adds the purchase path: units are added and removed in the API Platform account after an order form is signed, and "Any units purchased will be active until the start of the next invoice period, and renew daily after that."

Checked

Anthropic API service tiers documentation

"Priority Tier capacity commitments are no longer available for purchase"; a commitment consisted of "A number of input tokens per minute, A number of output tokens per minute, A commitment duration (1, 3, 6, or 12 months), A specific model version"; "Requests beyond your committed capacity automatically fall back to standard tier"; cache reads count as 0.1 tokens per token read.

Checked

Together AI pricing

"Reserve dedicated capacity in throughput units (PTUs). Each PTU represents fixed capacity. The tokens-per-minute it delivers depends on the model and the token type."; the PTU table's columns are Input TPM/PTU, Cached TPM/PTU, Output TPM/PTU, and Price PTU/MIN, but its per-model rows render client-side and did not load in our fetch, so no PTU figure is quoted here. Hardware is priced per GPU per hour: NVIDIA HGX H100 at $5.49 on-demand, HGX B200 at $8.99, and every reserved-hardware cell reading "Contact sales". Read again on 2026-09-29: the $5.49 list rate was unchanged; a dated promotional rate that ended on 09/30/26 is not carried in the article.

Checked

Fireworks AI pricing

On-demand deployments "Pay per GPU second, with no extra charges for start-up times". The hourly table carries two dated columns headed "Price ($) per hour - upto Aug 31" and "Price ($) per hour - from Sep 1", with no year stated; the current column reads H100 80 GB $8.00, H200 141 GB $8.00, B200 180 GB $13.00, B300 288 GB $15.00, and GB300 288 GB $20.00. Region-restricted deployments carry a 1.5x premium. No token-throughput entitlement is published for these deployments.

Checked

Baseten pricing

Dedicated deployments: "Only pay for the compute you use, down to the minute" and "you do not pay for idle time". Per-minute GPU instance rates run from T4 at $0.01052 to H100 80 GiB at $0.10833 and B200 180 GiB at $0.16633, with the entry plan at $0 per month, pay as you go, and volume discounts negotiated on higher plans. "You can deploy open source and custom models on Baseten."

Checked

Standard Thinking Product Plan

The unit definition (one model's base allocation of aggregate output tokens per second: Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300), $1,200 per unit per model per month, the $3,000 three-model bundle, extra capacity above the base that is free while the shared pool has room and is not reserved, auto scaling at "$5 per extra unit, per two-hour period" after authorization and subject to availability, no monthly token cap, and the statement that "The throughput figures do not describe a single request's speed or a monthly token allowance."

Checked

View all resources