Capacity and plan terms, defined

A working glossary for capacity plans and coding plans: 27 terms, each quoted from the vendor's own page or from ours, with the ones we have not published marked.

By Standard Thinking · Updated

Capacity plans and coding plans are sold in units that sound interchangeable and are not: a PTU, a Model Unit, a GSU, a token unit, a credit, and an aggregate output rate each count something different. This page defines the 27 terms you meet while comparing them, grouped into capacity units, plan mechanics, coding plan windows, and reading a rate. Third-party definitions quote the vendor's own page on the date we read it; ours repeat our published pages, and say where we have published none.

Capacity units

Provisioned throughput. Capacity bought ahead of the traffic that uses it. Google calls its version "a fixed-cost, fixed-term subscription ... that reserves throughput"; Azure bills a provisioned deployment hourly per PTU "regardless of the number of tokens consumed". Utilization is the buyer's risk; how LLM providers sell reserved capacity compares nine of them on their terms.

Provisioned throughput unit (PTU). Azure's billing object for a provisioned deployment. The PTU is common across models, but "the tokens per minute (TPM) that a given number of PTUs delivers depends on the model", and each model carries a minimum PTU count. Together AI sells a unit of the same name priced per PTU per minute, with separate input, cached, and output figures.

Model Unit (MU). Amazon Bedrock's unit, specifying "the number of input tokens that an MU can process across all requests within a span of one minute" and the matching output number. Both the specification and the price are private: the guide sends you to your AWS account manager for each.

Generative AI scale unit (GSU). Google Cloud's unit, "a measure of throughput for your prompts and responses". Per-model burndown rates convert input and output into that single figure, and per-second throughput means "your prompt input and generated output across all requests per second". Unused throughput does not carry over to the next month.

Token unit (OpenAI Scale Tier). Scale Tier sells "a set number of API input and output tokens per minute (known as 'token units') upfront for access to one specific model snapshot", each "purchased for a minimum of 30 days". Older models sell input and output units separately; GPT-5.4 and GPT-5.5 sell a combined bundle, and billing averages tokens per minute "in 15-minute intervals aligned to the top of the hour".

Tokens per minute (TPM). The usual way to state a reserved rate: a ceiling across all your requests inside a minute, not one request's speed. Azure, Bedrock, Together AI, and Anthropic's retired Priority Tier all use it. Two TPM figures compare when the model, the token type, and the measurement rules match, and not otherwise.

Aggregate output tokens per second. The Product Plan unit: output tokens produced across all of your product's requests together, for one model, with input work not counted in it. "The throughput figures do not describe a single request's speed or a monthly token allowance." What the unit means for your product turns the rate into answers per hour.

Unit (Product Plan). "One unit is that model's base allocation" of that rate: Kimi K3 at 40, GLM-5.3 at 100, and DeepSeek V4.1 Flash at 300 tokens per second, at $1,200 per unit, per model, per month. Units of the same model "scale that base in $1,200 steps", one allocation of each of the three is $3,000, and there is no monthly token cap.

Dedicated endpoint and GPU-hour. Renting the machine instead of a rate. Together AI lists an NVIDIA HGX H100 at $5.49 per GPU-hour and Fireworks AI bills on-demand deployments per GPU second, neither with a tokens-per-minute entitlement. The hardware is the ceiling, and throughput is whatever your model and traffic get out of it.

Reservation, and why it is not capacity. Usually a billing commitment rather than a hold on hardware. Azure states it plainly: "Reservations don't guarantee capacity. First create deployments to confirm that capacity is available, then purchase the reservation to lock in the discounted rate."

Spillover. What a provider does with requests above the reserved rate. Azure redirects them "to the standard deployment" on non-200 responses, though "Foundry Models from other providers (Azure DeepSeek, Meta Llama) don't currently support spillover".

Plan mechanics

Base allocation. The part of a plan reserved for the whole month, sized for the traffic you have most of the time. "Each model has a guaranteed base output rate shared across your product's concurrent requests", and "Your order defines the model, guaranteed base allocation, measurement conditions, and any service commitments." Size it from your busiest ordinary hour with the capacity and break-even calculator.

Auto scaling block. The two-hour period a peak is billed in: "$5 per extra unit, per two-hour period", added "after your authorization, subject to availability and plan terms". Blocks are counted rather than hours, so one hour of peak costs what two hours cost. Past 240 blocks in a month, a second base unit at $1,200 holds the same capacity for less.

Free extra capacity. Capacity above your base, "only while the shared pool has room. It is not reserved." It is not something to plan around: our calculator prices every hour above the base as auto scaling.

Use of the base. How much of the reserved rate a month's traffic consumes: the output the base itself served, each hour counted up to the base's rate, divided by what the base could have produced all month. Below 100% part of the rate goes unused; traffic above the base is not counted, because it draws on extra capacity rather than on what was reserved.

Break-even utilization. The level at which one unit, kept busy every second of the month at your input-to-output ratio and cached share, would cost the same paid per token as its monthly price. At 4:1 input to output and 0% cached against our published rates, that is 46% for GLM-5.3, 50% for Kimi K3, and 80% for DeepSeek V4.1 Flash: calculation assumptions, not measurements.

Input-to-output ratio. How many input tokens accompany each output token in a typical call. It moves the break-even more than a model's headline rate does: more input per output token pushes it down, and an agent that resends a transcript every turn sits at the high end. The chat and agent token estimator derives the ratio from your own call pattern.

Cached input and cache read. Input the provider has already processed, charged at a lower rate, moving the input side of a bill and nothing else. Anthropic's retired Priority Tier counted "cache reads as 0.1 tokens per token read from the cache"; Google expresses the same discount as a reduced burndown rate. Our pricing page publishes a cache-read rate per model; the matching rules and the usage field that reports a hit are not yet documented.

Pay per token. The baseline every capacity offer is measured against: input, cached-input, and output tokens charged as they are used, with no term and no reserved rate. It is the cheaper side while a unit would sit idle, and the dearer side once traffic is flat and busy. What an AI feature costs per user runs the same comparison per user.

Coding plan windows

5-hour window. A limit on a burst rather than on a month. Z.ai's "credit quota resets 5 hours after consumption", Kimi Code documents "a rolling 5-hour rate window", and Claude's session limit "will reset every five hours". Our $50 plan carries "5-hour and weekly limits" while the $100 and $200 plans have no 5-hour limit, and what the window counts is confirmed before activation.

Weekly window. The longer coding-plan limit. Z.ai's weekly credits are "Activated upon subscription; resets every 7 days", Kimi Code's legacy plans keep a 7-day quota while "for new members, the weekly quota limit is removed, and only the rolling 5-hour rate window remains", and Claude's "resets at a fixed time each week that is assigned to your account". Our $50 and $100 plans carry weekly limits; the $200 plan has "No 5-hour or weekly limits, and no total usage cap".

Rolling and fixed resets. Two clocks behind one window name. A rolling window returns the allowance a set time after it was spent, so waiting restores it in pieces; a fixed window refills at a single moment tied to a subscription date or to a time assigned to the account. Our reset rule is not published and is confirmed before activation.

Credits and model multipliers. An accounting unit that converts tokens into allowance, readable when the conversion is published. Z.ai publishes its formula — "(Input tokens × Input multiplier + Cached Input tokens × Cached Input multiplier + Output tokens × Output multiplier) / 10,000" — with GLM-5.3 at 6.9 input, 1.7 cached input, and 24 output, so a lighter model draws less allowance for the same work. Where none is published, an allowance cannot be converted into tokens; how coding plans count usage sets the units beside each other. Our own allowance unit is confirmed before activation.

Peak-time restriction and priority. "Peak-time restrictions apply to every plan, including $200. Requests may slow, queue, or time out. Priority is $200 → $100 → $50." Our pricing page adds that the order "is not a latency or availability guarantee", and which hours count as peak is not published. Z.ai publishes both parts: "During off-peak hours, model usage is charged at 50% of the standard credit rate. Peak hours: Monday to Friday, 14:00–18:00 Singapore Standard Time (UTC+8)."

Reading a rate

Answers per hour. A rate becomes a workload number once you fix an answer length: 3,600 × the rate ÷ the average output tokens in an answer. At 100 tokens per second and 400-token answers, that is 900 answers an hour, a ceiling at full use rather than a forecast.

Busiest normal hour. The traffic figure a base is sized from: the answers produced in an ordinary busy hour, not the record hour and not a daily average. The record hour belongs in the peak calculation, where capacity is bought in blocks.

What a throughput figure is not. "The throughput figures do not describe a single request's speed or a monthly token allowance." It is not a count of simultaneous users either, because converting a rate into people needs assumptions about sessions and overlap that the unit does not contain. Input throughput, concurrency, burst length, and latency targets are confirmed in your order rather than published.

Third-party definitions are quoted from the pages listed under Sources and verification; ours repeat our published pages. The API and both plans are open — log in to start.

Sources and verification

Sources used for this resource. Verification dates describe our source checks, not provider effective dates.

Azure AI Foundry provisioned throughput documentation

Provisioned deployments are billed "at an hourly rate ($/PTU/hr) based on the number of PTUs deployed, regardless of the number of tokens consumed"; "Reservations don't guarantee capacity"; "the tokens per minute (TPM) that a given number of PTUs delivers depends on the model" and each model has a minimum PTU count for a deployment; spillover "redirects those requests to the standard deployment" when a provisioned deployment returns non-200 responses, and "Foundry Models from other providers (Azure DeepSeek, Meta Llama) don't currently support spillover".

Checked

Amazon Bedrock Provisioned Throughput user guide

A Model Unit specifies "the number of input tokens that an MU can process across all requests within a span of one minute" and the matching number of output tokens it can generate; for "what an MU specifies, pricing per MU, and to request limit increases" the guide directs the reader to an AWS account manager rather than publishing either figure.

Checked

Vertex AI Provisioned Throughput overview

Provisioned Throughput is described as "a fixed-cost, fixed-term subscription available in several term-lengths that reserves throughput for supported generative AI models", which is the clearest one-sentence definition of buying capacity ahead of use among the pages surveyed.

Checked

Vertex AI, calculate Provisioned Throughput requirements

"A Generative AI Scale Unit (GSU) is a measure of throughput for your prompts and responses"; a burndown rate is "a ratio that converts the input and output units (such as tokens, characters, or images) to input tokens per second"; "Unused throughput doesn't accumulate or carry over to the next month".

Checked

Vertex AI, models supported by Provisioned Throughput

"Your per-second throughput is defined as your prompt input and generated output across all requests per second", with throughput, purchase increments and burndown rates published per model; our 2026-09-16 research records the Gemini Flash rows at 675 tokens per second per GSU, an output text token burning down 5 tokens and a cached input text token 0.1.

Checked

OpenAI Scale Tier for API customers

"Scale Tier lets you purchase a set number of API input and output tokens per minute (known as 'token units') upfront for access to one specific model snapshot. Each token unit is purchased for a minimum of 30 days." GPT-5.4 and GPT-5.5 sell a combined input and output bundle that "removes the need to predict your input and output token ratio"; tokens per minute are averaged "in 15-minute intervals aligned to the top of the hour", and tokens beyond the entitlement are "billed at pay-as-you-go (PAYG) rates on Fast mode".

Checked

Anthropic API service tiers documentation

"Priority Tier capacity commitments are no longer available for purchase." A commitment consisted of a number of input tokens per minute, a number of output tokens per minute, a duration of 1, 3, 6, or 12 months, and a specific model version, and counted "cache reads as 0.1 tokens per token read from the cache".

Checked

Together AI pricing

"Reserve dedicated capacity in throughput units (PTUs). Each PTU represents fixed capacity. The tokens-per-minute it delivers depends on the model and the token type." The PTU table carries separate Input, Cached and Output TPM columns and a price per PTU per minute; the hardware table lists NVIDIA HGX H100 at $5.49 per GPU-hour. Read again on 2026-09-29: the $5.49 list rate was unchanged; a dated promotional rate that ended on 09/30/26 is not carried.

Checked

Fireworks AI pricing

On-demand deployments are billed per GPU second, with start-up time not charged, and no tokens-per-minute or tokens-per-second entitlement is published for them: the deployed hardware is the ceiling.

Checked

Z.ai GLM Coding Plan documentation, Overview

The published credit formula, "Model credit usage = (Input tokens × Input multiplier + Cached Input tokens × Cached Input multiplier + Output tokens × Output multiplier) / 10,000", with GLM-5.3 at 6.9 input, 1.7 cached input and 24 output; "During off-peak hours, model usage is charged at 50% of the standard credit rate. Peak hours: Monday to Friday, 14:00-18:00 Singapore Standard Time (UTC+8)." Our 2026-09-12 research of the same documentation records the two reset rules: "5-hour credits: Dynamically refreshed; credit quota resets 5 hours after consumption" and "Weekly credits: Activated upon subscription; resets every 7 days".

Checked

Kimi Code documentation, Membership Benefits

Read again on 2026-09-29: "For new members, the weekly quota limit is removed, and only the rolling 5-hour rate window remains"; "Existing subscribers on legacy plans are not affected: the tier name, quota rules, and auto-renewal all remain as they are". On 2026-09-12 the same page described one 7-day quota ("refreshes automatically every 7 days from your subscription date") and the 5-hour window, with no new or legacy split.

Checked

Claude Help Center, What is the Max plan?

Recorded as the reference grammar for the five-hour and weekly pattern: "Your session-based usage limit will reset every five hours", and "Max plans also have a weekly usage limit that applies across all models. The weekly limit resets at a fixed time each week that is assigned to your account." The allowance itself is never published as an absolute unit, always as a multiple of another plan's usage per session.

Checked

Standard Thinking Product Plan

"Each model has a guaranteed base output rate shared across your product's concurrent requests" (Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300 tokens/s); "One unit is that model's base allocation"; $1,200 per unit, per model, per month with units of the same model added "to scale that base in $1,200 steps"; the three-model bundle at $3,000; "Capacity above your base only while the shared pool has room. It is not reserved."; "$5 per extra unit, per two-hour period" added "after your authorization, subject to availability and plan terms"; "Your order defines the model, guaranteed base allocation, measurement conditions, and any service commitments. The throughput figures do not describe a single request's speed or a monthly token allowance."

Checked

Standard Thinking pricing

A cache-read rate is published per model beside the input and output rates, in USD per 1M tokens (re-read on September 30, 2026); "A unit is one model's base allocation of output throughput"; and for the coding tiers, "Priority is $200 to $100 to $50; it is not a latency or availability guarantee" and "Allowance amounts, supported tools, renewal dates, and cancellation terms are confirmed before activation."

Checked

Standard Thinking Coding Plan

The $50 tier carries a smaller allowance with "5-hour + weekly limits", the $100 tier a larger allowance with "Weekly limit only", and the $200 tier "No 5-hour or weekly limits, and no total usage cap"; "Peak-time restrictions apply to every plan: requests may slow, queue, or time out. Priority is $200 to $100 to $50." The allowance unit and the reset rule are not stated on the page.

Checked

Standard Thinking documentation

The Coding Plan limits table gives the $50 tier "5-hour and weekly limits" and the $200 tier "No 5-hour or weekly limits"; "Peak-time restrictions apply to every plan, including $200. Requests may slow, queue, or time out."; "Exact allowances and request and concurrency limits are confirmed before activation." The chat completions response documents usage.prompt_tokens and usage.completion_tokens and no field for cached tokens.

Checked

View all resources