What “aggregate output tokens per second” means for your product
A throughput unit is a rate shared by every request at once. Read it as answers per hour, and size peaks in 2-hour blocks.
Aggregate output tokens per second is the rate at which one model unit produces output tokens across all of your product's requests together. It is a rate, not a quantity: it does not describe how quickly a single request finishes, and it is not a monthly token allowance. When ten answers are being generated at once, they share the unit's rate. Product Plan sells that rate as a monthly plan, at a published price per model per month and with no monthly token cap, so the planning question is not “how many tokens will I use this month” but “what output rate does my busiest normal hour need”.
From a rate to a number of answers
A rate becomes a workload number once you decide how long an answer is. With R as the unit's output rate and T as the average output tokens in one answer:
answers per hour = 3,600 × R ÷ T
answers per day = 24 × answers per hour
The figures below are a calculation assumption, not a measurement: a 400-token average answer, a 30-day month, and a unit fully used every second of the day. No product sustains full use, so read them as a unit's ceiling, not a forecast.
| Unit output rate | Answers per hour | Answers per day | Output tokens per 30-day month |
|---|---|---|---|
| 40 tokens / s (Kimi K3) | 360 | 8,640 | 103.68M |
| 100 tokens / s (GLM-5.3) | 900 | 21,600 | 259.2M |
| 300 tokens / s (DeepSeek V4.1 Flash) | 2,700 | 64,800 | 777.6M |
Answer length moves these numbers in inverse proportion: at 100 tokens per second, 800-token answers give 450 per hour instead of 900. How busy the unit is moves the day column: a unit busy half the time produces half as many answers, for the same monthly price.
Why input tokens are not in the unit
The unit counts output tokens only. Reading the prompt, the conversation history, and any tool results is input work, and it is not what 40, 100, or 300 measures. Output generation is also the part of a response that streams to your users.
It also means the unit does not answer every capacity question by itself. Input throughput, how many requests may be in flight at once, how long a burst may run, and any latency target are confirmed in your order rather than published. The Product Plan page states the boundary directly: your order defines the model, the guaranteed base allocation, the measurement conditions, and any service commitments, and the throughput figures do not describe a single request's speed or a monthly token allowance.
For the same reason, do not turn a rate into a headcount. Converting a rate into people needs assumptions about session length, think time, and overlapping demand that the unit does not contain. Answers per hour is a conversion the unit supports; people are not.
Peaks: extra units or 2-hour blocks
A base unit runs for the whole month, so it is sized for the traffic you have most of the time, not for the worst hour of the month. While the shared pool has room, extra capacity above your base is free, but it is not reserved and not something to plan around. When you need the extra rate to be there, auto scaling adds units at $5 per extra unit per 2-hour period, after your authorization and subject to available capacity.
A peak becomes a small, countable cost. A six-hour peak needing one extra unit spans three 2-hour blocks and costs $15; the same peak on about 22 weekdays is 66 blocks, or $330. Both are calculation assumptions applied to the published $5 block price.
The crossover with a second base unit is arithmetic: $1,200 ÷ $5 = 240 blocks. Below 240 blocks of the same extra unit in a month, auto scaling costs less. Past 240 blocks — 480 hours, about two-thirds of a 30-day month — a second base unit at $1,200 costs less, and it is reserved.
How other providers express the same idea
Reserved capacity is sold in throughput units nearly everywhere; what differs is how the unit is described.
- Microsoft's Azure AI Foundry documentation sells provisioned throughput in PTUs and expresses the throughput as tokens per minute, with the tokens per minute per PTU varying by model. It also notes that cached tokens do not consume PTU capacity.
- Amazon's Bedrock Provisioned Throughput guide defines a Model Unit by the input and output tokens per minute it can process across all requests, and bills Provisioned Throughput hourly.
Other providers sell similar units. These numbers do not convert into one another: they count different token types (input and output together, output only, cached separately), they are tied to specific models, and each is measured under its own rules. A throughput number only means something with the model, the token type, and the measurement conditions attached.
Turn your traffic into a rate
Four steps, from numbers you have:
- Count the answers your product produced in its busiest normal hour — not the record hour, and not a daily average.
- Multiply by the average output tokens in one answer. If you do not have that number, estimate it with the chat and agent token estimator, which accounts for conversation history and repeated tool calls.
- Divide by 3,600. That is the aggregate output rate your busiest normal hour needs, in tokens per second.
- Compare it with 40, 100, and 300, and round up to whole units. Then list the hours that exceed your base and count them in 2-hour blocks.
The Product Plan capacity and break-even calculator runs these four steps from your traffic and adds the pay-per-token comparison beside them.
A worked example, as a calculation assumption: a product answering 600 questions in its busiest normal hour at 500 output tokens each needs 600 × 500 ÷ 3,600 = about 83 tokens per second, inside one GLM-5.3 unit at 100. A campaign that doubles that hour for four hours twice a week is two blocks per event and about eight events a month: 16 blocks, or $80 of auto scaling.
Sizing is half the decision. Whether a unit costs less than paying per token depends on your input-to-output ratio and cached share more than on the output rate itself; how to buy AI inference for a product on a monthly plan works through that comparison.
Product Plan is available now: log in to order a base, or discuss your traffic with us first.
Sources and verification
Sources used for this resource. Verification dates describe our source checks, not provider effective dates.
Standard Thinking Product Plan
The unit definition (one unit is one model's base allocation of aggregate output tokens per second: Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300), $1,200 per unit per model per month, auto scaling at $5 per extra unit per 2-hour period after authorization and subject to available capacity, free extra capacity while the shared pool has room, no monthly token cap, and the sentence that the order defines the model, guaranteed base allocation, measurement conditions and any service commitments while the throughput figures do not describe a single request's speed or a monthly token allowance.
Microsoft Azure AI Foundry provisioned throughput documentation
Provisioned throughput is sold in PTUs and the throughput per PTU is expressed as tokens per minute, which varies by model; cached tokens do not consume PTU capacity.
Amazon Bedrock Provisioned Throughput user guide
A Model Unit specifies the input and output tokens per minute it can process across all requests, and Provisioned Throughput is billed hourly.
Continue exploring
$3,000a month for the example product
- Pay per token
- $4,500
- Base
- 2 GLM-5.3 units
- Auto scaling
- 60 blocks
Calculation example, not a measurement
How many Product Plan units does your traffic need, and what would pay per token cost?
Size a monthly plan from your busiest ordinary hour, price the peaks in blocks, and read the month against pay per token.
46%utilization where API cost equals one unit
- Model
- GLM-5.3
- Unit tok/s
- 100
- API total
- $2,592.00
Calculation example, not a measurement
How to buy AI inference for a product on a monthly plan
The four ways a product can pay for inference, what each one actually promises, and a worked break-even against our published rates.
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.