Compare prices and plan limits.
Pay per token with Standard One or API, subscribe to Product Plan for your product, or choose Coding Plan for your development work.
Standard One
Choose a size and pay per input token. Each call returns one short answer, so output is free.
Estimate your monthly Standard One cost.
$2.26 / month
- Text input
- $1.90
- Image input
- $0.36
- Output
- $0.00
- Total token volume
- 110M
Planning estimate only. Actual charges follow metered usage and exclude tax and separately agreed services.
Standard API
Compare input, output, and cache-read rates for each model.
| Model | Input / 1M tokens | Output / 1M tokens | Cache read / 1M tokens | Context window |
|---|---|---|---|---|
| 1.05M | ||||
| our price $1.40 | our price $4.40 | our price $0.26 | 1.05M | |
| 1.31M | ||||
| our price $0.037 | our price $0.17 | — | 131K | |
| our price $0.25 | 1.05M | |||
| 1.05M | ||||
| 1.05M | ||||
| 262K |
Estimate your monthly API cost.
Estimate a chat or agent workload$4.56 / month
- Input
- $1.68
- Output
- $2.88
- Total token volume
- 10M
Planning estimate only. Actual charges follow metered usage and exclude tax and separately agreed services.
Product Plan
A unit is one model’s base allocation of output throughput. Choose how many units you keep all month, and add more for busy periods in 2-hour blocks.
Monthly base
$1,200 per unit, per model, per month
Reserved capacity on one model, available at every hour of the month.
Example: two units of GLM-5.3 are $2,400 a month.
Auto scaling
$5 per extra unit, per 2-hour period
Extra units you add on top of your base during a traffic peak, billed only for the 2-hour periods they run.
Example: one extra unit for a 6-hour peak is $15.
Estimate your monthly Product Plan cost.
$1,205 / month
- Monthly base 1 × $1,200
- $1,200
- This peak 1 extra unit × 1 period × $5
- $5
Planning estimate, before tax. Extra peak periods in the month are charged separately. For an estimate from your own traffic, use the Product Plan capacity calculator.
What one busy day looks like.
Your base runs all month. Each peak block is one 2-hour period of auto scaling on top of it.
Monthly baseAuto scaling
Auto scaling requires your authorization and available capacity. Billing is monthly with a one-month minimum; adding units mid-month charges only the prorated difference for the rest of the month.
| Charge | Price | What it covers |
|---|---|---|
| Kimi K3 allocation | $1,200 per unit, per model, per month | 40 aggregate output tokens per second. |
| GLM-5.3 allocation | $1,200 per unit, per model, per month | 100 aggregate output tokens per second. |
| DeepSeek V4.1 Flash allocation | $1,200 per unit, per model, per month | 300 aggregate output tokens per second. |
| Three-model bundle | $3,000 per month | One allocation each of Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Save $600 per month compared with three individual plans. |
| Free extra capacity | No extra charge | Capacity above your base only while the shared pool has room. It is not reserved. |
| Auto scaling | $5 per extra unit, per two-hour period | Auto scaling adds capacity for traffic peaks after your authorization, subject to availability and plan terms. |
Coding Plan
All plans include Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Compare allowance timing and busy-period priority as well as price.
$50 per developer, per month
- Smaller allowance
- 5-hour + weekly limits
- Peak-time restrictions
- Lowest priority
$100 per developer, per month
- Larger allowance
- Weekly limit only, no 5-hour limit
- Peak-time restrictions
- Priority above $50, below $200
$200 Highest priority per developer, per month
- No 5-hour or weekly limit
- No total usage cap
- Peak-time restrictions
- Highest priority
Allowance amounts, supported tools, renewal dates, and cancellation terms are confirmed before activation. Peak-time restrictions apply to every tier, including $200: requests may slow, queue, or time out. Priority is $200 → $100 → $50; it is not a latency or availability guarantee.
Need help comparing the options?
Tell us whether you are building an application or choosing a coding plan.
API or Product Plan?
Estimate API cost from input and output tokens. For a Product Plan, choose a model allocation and account for auto scaling charges. Compare service conditions as well as the total. How to buy AI inference for a product on a monthly plan walks through the comparison.
What does Coding Plan include?
A monthly plan for one developer with Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Allowance timing and busy-period priority differ by tier. Check access and setup
Does an account activate a plan?
Account creation, product access, and plan activation are separate steps. Check access and plan conditions before activation.