Product Plan · GLM-5.3
How many Product Plan units does your traffic need, and what would pay per token cost?
Size a monthly plan from your busiest ordinary hour, price the peaks in blocks, and read the month against pay per token.
Plan terms and rates checked . One unit is one model's base allocation of aggregate output tokens per second — GLM-5.3 at 100, Kimi K3 at 40, DeepSeek V4.1 Flash at 300 — at $1,200 per unit, per model, per month, with auto scaling at $5 per extra unit per 2-hour block.
Size a plan for your traffic
Standard API · pay per token
Normal-hour rate
What an ordinary hour asks of the plan
138.9 tokens / s
Peak-hour rate
What the busiest hour asks
347.2 tokens / s
Use of the base
Output the base itself served, against its capacity
75%
Break-even use
Where one unit's pay-per-token cost reaches its monthly price
46%
Answers an hour
What the base produces at this answer length, fully used
1,800
Plan terms and rates checked . This is a dated snapshot of published terms; confirm current terms before budgeting. Every figure here is a calculation from your assumptions, not a measurement.
Plan terms and published rates
| Model | One unit | Input | Cache read | Output | Source |
|---|---|---|---|---|---|
| GLM-5.3Unit source | 100 | $1.40 | $0.26 | $4.40 | Rate source |
| Kimi K3Unit source | 40 | $2.55 | $0.255 | $12.75 | Rate source |
| DeepSeek V4.1 FlashUnit source | 300 | $0.24 | $0.0048 | $0.96 | Rate source |
Unit rates in aggregate output tokens per second · token rates in USD per 1 million tokens · $1,200 per unit, per model, per month · $5 per extra unit per 2-hour block
How to read the result
The calculator asks for a traffic pattern and answers two questions at once: how much reserved capacity that traffic needs, and what the same traffic would cost paid per token. Four readings carry most of the meaning.
A plan is sized by rate, not by volume. The answers in an ordinary hour, multiplied by the tokens in an answer and divided by the seconds in an hour, give the aggregate output rate that hour asks for. That rate is compared with one unit's output rate and rounded up, because capacity is sold in whole units and a plan starts at one. Everything on the plan side follows from that number: the base cost is the units multiplied by the unit's monthly price, whatever the month's token total turns out to be.
Peaks are bought by the block, not by the month. The busiest hour is sized the same way, and the difference between the units it needs and the base is what auto scaling has to supply. Blocks are counted, not hours: the peak hours in a day are rounded up to whole blocks and multiplied by the number of days the peak repeats. Two hours of peak and one hour of peak cost the same in a day, so a short, sharp peak is the cheap kind.
The pay-per-token column is the same traffic, priced per token. Normal hours and peak hours are added into one month of answers, multiplied by the input and output tokens an answer, and charged at the published rates in the table under Plan terms and published rates. The cached share moves only the input line; it never changes the plan price.
The two utilization figures answer different questions. Use of the base is how much of the reserved rate this traffic actually consumes over the month: capacity sits idle below 100%, and above the base the traffic is drawing on extra capacity rather than on what was reserved, so the figure stops there. Break-even use is the level at which one unit, used every second of the month at your input-to-output ratio and cached share, would cost the same paid per token as the unit's monthly price. Traffic that keeps a unit busier than its break-even level costs less on the plan; traffic that leaves it quieter costs less paid per token. The ratio and the cached share move that line more than the model's headline rate does, which is why both are inputs.
The calculation
normal rate = answers in a normal hour × output tokens ÷ 3,600
base units = max(1, ceil(normal rate ÷ one unit's output rate))
base cost = base units × the unit's monthly price
peak rate = answers in a peak hour × output tokens ÷ 3,600
units at peak = max(1, ceil(peak rate ÷ one unit's output rate))
extra units = units at peak − base units
blocks = ceil(peak hours a day ÷ hours in a block) × peak days a month
auto scaling = extra units × blocks × the price of a block
answers a month = normal answers × (hours in the month − peak hours)
+ peak answers × peak hours
pay per token = (uncached input × input rate + cached input × cache-read rate
+ output × output rate) ÷ 1,000,000
base rate = base units × unit rate
use of the base = (normal hours × min(normal rate, base rate)
+ peak hours × min(peak rate, base rate)) × 3,600
÷ (base rate × seconds in the month)
break-even use = the unit's monthly price ÷ the pay-per-token cost of one unit used every second
answers an hour = 3,600 × unit rate × base units ÷ output tokens
Calculation uses unrounded values; only the displayed figures are rounded. Reset restores the example traffic, which is an assumption for reading the page rather than observed customer usage. Every figure here is a calculation from the inputs above and the published terms in the table, not a measurement of any product.
What this calculator does not decide
A unit is defined by aggregate output tokens per second, so that is all this arithmetic can size. Input throughput, how many requests may be in flight at once, how long a burst may run, and any latency target are confirmed in your order rather than published, and no figure here stands in for them. For the same reason the result is never a number of simultaneous users: converting a rate into people needs assumptions about sessions and overlap that the unit does not contain.
Two plan terms are deliberately left out of the arithmetic. Extra capacity above your base is free while the shared pool has room, but it is not reserved, so counting on it would flatter the estimate; the calculator prices every hour above the base as auto scaling. The three-model bundle is not modelled either — the comparison here is one model at a time, against the same model's token rates.
The pay-per-token column covers input, cached-input, and output token charges only. It excludes taxes, credits, negotiated terms, tool execution, storage, and networking. And a plan, once ordered, is a monthly commitment: a product whose traffic collapses mid-month still holds the capacity it reserved.
From an estimate to an order
Start with the two numbers the whole result turns on. Count the answers your product produced in its busiest ordinary hour — not the record hour, and not a daily average — and measure the average output tokens in an answer; the chat and agent token estimator will give you the second one from your own call pattern, including conversation history and repeated tool calls. Then list the hours that genuinely exceed that baseline and count them in blocks.
If the result is close either way, read how to buy AI inference for a product on a monthly plan for what each buying model actually promises, and what "aggregate output tokens per second" means for your product for the rate arithmetic in more detail. Bring your measured hour and answer length to the conversation; an order is written around measurement conditions, not around an estimate.
Product Plan is available now: log in to order units, or discuss your workload with us with your measured hour and answer length. The plan terms and rates above are our published terms; everything computed from them is a calculation from your assumptions.
Sources and verification
Sources used for this resource. Verification dates describe our source checks, not provider effective dates.
Standard Thinking Product Plan
One unit is one model's guaranteed base allocation of aggregate output tokens per second: Kimi K3 40, GLM-5.3 100, DeepSeek V4.1 Flash 300. A unit is $1,200 per model per month and units of the same model add in $1,200 steps; there is no monthly token cap. Auto scaling adds capacity at $5 per extra unit per 2-hour period, after your authorization and subject to available capacity, and extra capacity above the base is free while the shared pool has room without being reserved. Your order defines the model, the guaranteed base allocation, the measurement conditions and any service commitments; the throughput figures do not describe a single request's speed or a monthly token allowance. Re-read on September 30, 2026: unit sizes, the $1,200 unit price, the $3,000 bundle and the $5 auto-scaling block were unchanged.
Standard API rates in USD per 1M tokens used for the pay-per-token side of this calculator, re-read on September 30, 2026: GLM-5.3 $1.40 input, $4.40 output, $0.26 cache read; Kimi K3 $2.55, $12.75, $0.255; DeepSeek V4.1 Flash $0.24, $0.96, $0.0048.
Continue exploring
46%utilization where API cost equals one unit
- Model
- GLM-5.3
- Unit tok/s
- 100
- API total
- $2,592.00
Calculation example, not a measurement
How to buy AI inference for a product on a monthly plan
The four ways a product can pay for inference, what each one actually promises, and a worked break-even against our published rates.
900answers per hour
- Throughput
- 100 tokens / s
- Answer length
- 400 tokens
- Utilization
- 100%
Calculation example, not a measurement
What “aggregate output tokens per second” means for your product
A throughput unit is a rate shared by every request at once. Read it as answers per hour, and size peaks in 2-hour blocks.
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.