What “aggregate output tokens per second” means for your product

A throughput unit is a rate shared by every request at once. Read it as answers per hour, and size peaks in 2-hour blocks.

By Standard Thinking · Updated

Aggregate output tokens per second is the rate at which one model unit produces output tokens across all of your product's requests together. It is a rate, not a quantity: it does not describe how quickly a single request finishes, and it is not a monthly token allowance. When ten answers are being generated at once, they share the unit's rate. Product Plan sells that rate as a monthly plan, at a published price per model per month and with no monthly token cap, so the planning question is not “how many tokens will I use this month” but “what output rate does my busiest normal hour need”.

From a rate to a number of answers

A rate becomes a workload number once you decide how long an answer is. With R as the unit's output rate and T as the average output tokens in one answer:

answers per hour = 3,600 × R ÷ T
answers per day  = 24 × answers per hour

The figures below are a calculation assumption, not a measurement: a 400-token average answer, a 30-day month, and a unit fully used every second of the day. No product sustains full use, so read them as a unit's ceiling, not a forecast.

Unit output rate Answers per hour Answers per day Output tokens per 30-day month
40 tokens / s (Kimi K3) 360 8,640 103.68M
100 tokens / s (GLM-5.3) 900 21,600 259.2M
300 tokens / s (DeepSeek V4.1 Flash) 2,700 64,800 777.6M

Answer length moves these numbers in inverse proportion: at 100 tokens per second, 800-token answers give 450 per hour instead of 900. How busy the unit is moves the day column: a unit busy half the time produces half as many answers, for the same monthly price.

Why input tokens are not in the unit

The unit counts output tokens only. Reading the prompt, the conversation history, and any tool results is input work, and it is not what 40, 100, or 300 measures. Output generation is also the part of a response that streams to your users.

It also means the unit does not answer every capacity question by itself. Input throughput, how many requests may be in flight at once, how long a burst may run, and any latency target are confirmed in your order rather than published. The Product Plan page states the boundary directly: your order defines the model, the guaranteed base allocation, the measurement conditions, and any service commitments, and the throughput figures do not describe a single request's speed or a monthly token allowance.

For the same reason, do not turn a rate into a headcount. Converting a rate into people needs assumptions about session length, think time, and overlapping demand that the unit does not contain. Answers per hour is a conversion the unit supports; people are not.

Peaks: extra units or 2-hour blocks

A base unit runs for the whole month, so it is sized for the traffic you have most of the time, not for the worst hour of the month. While the shared pool has room, extra capacity above your base is free, but it is not reserved and not something to plan around. When you need the extra rate to be there, auto scaling adds units at $5 per extra unit per 2-hour period, after your authorization and subject to available capacity.

A peak becomes a small, countable cost. A six-hour peak needing one extra unit spans three 2-hour blocks and costs $15; the same peak on about 22 weekdays is 66 blocks, or $330. Both are calculation assumptions applied to the published $5 block price.

The crossover with a second base unit is arithmetic: $1,200 ÷ $5 = 240 blocks. Below 240 blocks of the same extra unit in a month, auto scaling costs less. Past 240 blocks — 480 hours, about two-thirds of a 30-day month — a second base unit at $1,200 costs less, and it is reserved.

How other providers express the same idea

Reserved capacity is sold in throughput units nearly everywhere; what differs is how the unit is described.

Other providers sell similar units. These numbers do not convert into one another: they count different token types (input and output together, output only, cached separately), they are tied to specific models, and each is measured under its own rules. A throughput number only means something with the model, the token type, and the measurement conditions attached.

Turn your traffic into a rate

Four steps, from numbers you have:

  1. Count the answers your product produced in its busiest normal hour — not the record hour, and not a daily average.
  2. Multiply by the average output tokens in one answer. If you do not have that number, estimate it with the chat and agent token estimator, which accounts for conversation history and repeated tool calls.
  3. Divide by 3,600. That is the aggregate output rate your busiest normal hour needs, in tokens per second.
  4. Compare it with 40, 100, and 300, and round up to whole units. Then list the hours that exceed your base and count them in 2-hour blocks.

The Product Plan capacity and break-even calculator runs these four steps from your traffic and adds the pay-per-token comparison beside them.

A worked example, as a calculation assumption: a product answering 600 questions in its busiest normal hour at 500 output tokens each needs 600 × 500 ÷ 3,600 = about 83 tokens per second, inside one GLM-5.3 unit at 100. A campaign that doubles that hour for four hours twice a week is two blocks per event and about eight events a month: 16 blocks, or $80 of auto scaling.

Sizing is half the decision. Whether a unit costs less than paying per token depends on your input-to-output ratio and cached share more than on the output rate itself; how to buy AI inference for a product on a monthly plan works through that comparison.

Product Plan is available now: log in to order a base, or discuss your traffic with us first.

Sources and verification

Sources used for this resource. Verification dates describe our source checks, not provider effective dates.

Standard Thinking Product Plan

The unit definition (one unit is one model's base allocation of aggregate output tokens per second: Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300), $1,200 per unit per model per month, auto scaling at $5 per extra unit per 2-hour period after authorization and subject to available capacity, free extra capacity while the shared pool has room, no monthly token cap, and the sentence that the order defines the model, guaranteed base allocation, measurement conditions and any service commitments while the throughput figures do not describe a single request's speed or a monthly token allowance.

Checked

Microsoft Azure AI Foundry provisioned throughput documentation

Provisioned throughput is sold in PTUs and the throughput per PTU is expressed as tokens per minute, which varies by model; cached tokens do not consume PTU capacity.

Checked

Amazon Bedrock Provisioned Throughput user guide

A Model Unit specifies the input and output tokens per minute it can process across all requests, and Provisioned Throughput is billed hourly.

Checked

View all resources