How to buy AI inference for a product on a monthly plan

The four ways a product can pay for inference, what each one actually promises, and a worked break-even against our published rates.

By Standard Thinking · Updated

A product with steady AI traffic can buy inference per token, by the GPU hour, by reserved throughput, or on a monthly plan, sold as a seat or a throughput unit. At what level of steady use does capacity cost less than tokens? Against our published rates, a $1,200 monthly unit breaks even at 46% utilization on GLM-5.3, 50% on Kimi K3, and 80% on DeepSeek V4.1 Flash.

Four ways to buy inference

Terms come from each vendor's own page, checked on September 16, 2026 and linked under Sources and verification. Nine offerings are set side by side in how LLM providers sell reserved capacity.

Way Unit Billing Minimum term Token cap
Pay per token Input, cached, output tokens Per request None None
Dedicated GPU hour One rented GPU instance Hourly None Hardware limit only
Provisioned throughput Tokens per minute Hourly None to 1 year None inside the unit
Flat-rate subscription Seat with concurrency limit Monthly 1 month None, concurrency capped instead
Monthly throughput plan Aggregate output tokens per second Monthly plus 2-hour blocks 1 month None

Notes on each way

  • Pay per token. Every input, cached-input, and output token is charged as it is used; the break-even below uses our published Standard API rates.
  • Dedicated GPU hour. Together AI lists an NVIDIA HGX H100 endpoint at $5.49 per GPU-hour.
  • Provisioned throughput. Azure PTUs and Bedrock Model Units bill hourly regardless of tokens consumed, on terms of none, 1 month, 6 months, or 1 year by vendor. Together AI prices its PTUs per minute; Bedrock's unit specification and price come from an AWS account manager.
  • Flat-rate subscription. Featherless sells a $25 Chat plan with 4 concurrent units and no token counting, then excludes "app or API traffic, reselling, background automation and benchmarking". Its Business plan is a custom contract sized to the buyer's GPUs.
  • Monthly throughput plan. Product Plan is $1,200 per model per month for one unit: Kimi K3 at 40, GLM-5.3 at 100, DeepSeek V4.1 Flash at 300 aggregate output tokens per second. One allocation of each is $3,000.

What a reserved hour actually buys

A reserved hour, GPU or throughput, is charged regardless of the tokens consumed, so utilization is the buyer's risk and the break-even is worth settling before a term starts. The unit is throughput, and it moves with the model: an Azure PTU is model-independent, but the tokens per minute it delivers and the minimum deployment size differ by model. Anthropic's retired Priority Tier had the same shape: input and output tokens per minute, a term of 1, 3, 6, or 12 months, a named model version, and a standard-tier fallback. It is "no longer available for purchase".

A reservation is a commitment, not a capacity promise

Azure states that "Reservations don't guarantee capacity": a reservation discounts the hours you buy rather than holding hardware. Azure's guidance also advises against scaling provisioned deployments with traffic, and spillover to a pay-per-token deployment covers Azure OpenAI models rather than other providers' models hosted there, so a peak above the reservation has to be planned for rather than absorbed.

Where Product Plan sits

Product Plan sells what the hourly reservations sell — throughput held for your product — as a monthly plan with a published price. Units of the same model add in $1,200 steps, there is no monthly token cap, and capacity above the base allocation is free while the shared pool has room, though it is not reserved. Your order settles the model, the base allocation, the measurement conditions, and input-side conditions such as input throughput, concurrency, burst length, and latency targets. The unit is an aggregate rate: not one request's speed, not a monthly token allowance, and not a count of simultaneous users.

When a monthly unit beats pay per token

The table below is a calculation assumption, not a measurement and not an offer: a 30-day month, a unit used fully for the whole month, 1,600 input and 400 output tokens per call at 4:1, 0% cached input, taxes excluded, and our published Standard API rates per 1M tokens — GLM-5.3 $1.40 and $4.40, Kimi K3 $2.55 and $12.75, DeepSeek V4.1 Flash $0.24 and $0.96 — as published on the pricing page and re-read on September 30, 2026.

Model Unit tok/s API total Break-even utilization
GLM-5.3 100 $2,592.00 46%
Kimi K3 40 $2,379.46 50%
DeepSeek V4.1 Flash 300 $1,492.99 80%

More input per output token pushes the break-even down; more cached input pushes it up. An agent workload that resends a long transcript sits below these figures, a short prompt with a long answer above them. A month of full use is 259.2M output tokens on GLM-5.3, 103.68M on Kimi K3, and 777.6M on Flash. The Flash unit is not a savings story at 80%: its case is holding 300 tokens per second at a fixed monthly price rather than paying less per token.

Tian Pan's post on reserved capacity (July 4, 2026) puts the break-even "somewhere around 60-80% sustained utilization" — a rule of thumb rather than a measurement, above our GLM-5.3 and Kimi K3 figures and below Flash.

What peaks cost

Auto scaling adds capacity at $5 per extra unit per 2-hour period, after your authorization and subject to available capacity. It beats a second base unit until a month needs more than 240 blocks, or 480 hours of that unit; past that point the $1,200 unit holds the same capacity for less. Run the arithmetic on your own traffic in the Product Plan capacity and break-even calculator.

Three questions to answer before choosing

  1. How flat is your baseline? Capacity pricing rewards steady traffic; demand that collapses overnight pays for the quiet hours unless batch work moves in.
  2. What is your input-to-output ratio, and how much is cached? The chat and agent token estimator gives both; caching rules decide what counts as cached.
  3. What is your peak multiple? Compare the busiest normal hour to the average, then read the gap as aggregate output tokens per second.

Product Plan and the Standard API are both available now: log in to start, or discuss your workload with us before choosing a base.

Sources and verification

Sources used for this resource. Verification dates describe our source checks, not provider effective dates.

Azure AI Foundry provisioned throughput documentation

Provisioned deployments are billed at an hourly rate ($/PTU/hr) regardless of the number of tokens consumed; Azure Reservations run for 1 month or 1 year and "Reservations don't guarantee capacity"; the PTU itself is model-independent while the tokens per minute each PTU delivers vary by model, with a minimum deployment size per model; the page advises against scaling provisioned deployments up and down with traffic; spillover to a standard deployment is available for Azure OpenAI models rather than models from other providers.

Checked

Amazon Bedrock Provisioned Throughput user guide

Provisioned Throughput is billed hourly; a Model Unit specifies input and output tokens per minute across all requests; commitment terms are none, 1 month, or 6 months; the Model Unit specification and its price come from your AWS account manager rather than a published rate card.

Checked

Together AI pricing

Provisioned Throughput Units are described as fixed capacity with Input, Cached, and Output tokens per minute per PTU and are priced per PTU per minute; dedicated NVIDIA HGX H100 endpoints are listed at $5.49 per GPU-hour and B200 at $8.99. Read again on 2026-09-29: the $5.49 list rate was unchanged; a dated promotional rate that ended on 09/30/26 is not carried.

Checked

Anthropic API service tiers documentation

"Priority Tier capacity commitments are no longer available for purchase"; a commitment combined input tokens per minute, output tokens per minute, a term of 1, 3, 6, or 12 months, and a specific model version, and requests beyond the commitment fell back to the standard tier.

Checked

Featherless AI pricing

The Chat plan is $25 per month with a 32K context and "4 concurrent units", and excludes "app or API traffic, reselling, background automation and benchmarking"; the Business plan is a custom contract "Sized to your GPUs".

Checked

Tian Pan, "Reserved capacity for tokens" (opinion, July 4, 2026)

A practitioner's heuristic in an opinion post, not a measurement: the author puts the break-even for reserved capacity "somewhere around 60-80% sustained utilization".

Checked

Standard Thinking Product Plan

The unit definition (one model's guaranteed base allocation of aggregate output tokens per second: Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300), $1,200 per unit per model per month, the $3,000 three-model bundle, free extra capacity while the shared pool has room, auto scaling at $5 per unit per 2-hour period after authorization, no monthly token cap, and the statement that the throughput figures do not describe a single request's speed or a monthly token allowance.

Checked

Standard Thinking pricing

The Standard API rates used in this page's calculation, in USD per 1M tokens: GLM-5.3 $1.40 input and $4.40 output, Kimi K3 $2.55 and $12.75, DeepSeek V4.1 Flash $0.24 and $0.96. Unchanged when re-read on September 30, 2026.

Checked

View all resources