Skip to content

Compare prices and plan limits.

Pay per token with Standard One or API, subscribe to Product Plan for your product, or choose Coding Plan for your development work.

Standard One

Choose a size and pay per input token. Each call returns one short answer, so output is free.

Model Text input / 1M tokens Image input / 1M tokens Output Model ID Specs
One 3B 50% off Faster, for high-volume calls. list price $0.038 our price $0.019 list price $0.072 our price $0.036 $0 standard-one-3b
One 8B 50% off Stronger on hard cases, for high-stakes calls. list price $0.098 our price $0.049 list price $0.168 our price $0.084 $0 standard-one-8b

Estimate your monthly Standard One cost.

Workload inputs

Estimated monthly Standard One cost

$2.26 / month

Text input
$1.90
Image input
$0.36
Output
$0.00
Total token volume
110M

Planning estimate only. Actual charges follow metered usage and exclude tax and separately agreed services.

Standard API

Compare input, output, and cache-read rates for each model.

Model Input / 1M tokens Output / 1M tokens Cache read / 1M tokens Context window
DeepSeek DeepSeek-V4.1-Flash 20% off list price $0.30 our price $0.24 list price $1.20 our price $0.96 list price $0.006 our price $0.0048 1.05M
Z.ai GLM-5.3 our price $1.40 our price $4.40 our price $0.26 1.05M
Z.ai GLM-5.3-Flash 15% off list price $0.15 our price $0.1275 list price $0.50 our price $0.425 list price $0.03 our price $0.0255 1.31M
OpenAI gpt-oss-120b our price $0.037 our price $0.17 — 131K
Qwen Qwen3.8-2.4T-A95B 20% off list price $2.00 our price $1.60 list price $6.00 our price $4.80 our price $0.25 1.05M
DeepSeek DeepSeek-V4-Pro-0813 20% off list price $1.32 our price $1.056 list price $3.96 our price $3.168 list price $0.044 our price $0.0352 1.05M
Moonshot AI Kimi K3 15% off list price $3.00 our price $2.55 list price $15.00 our price $12.75 list price $0.30 our price $0.255 1.05M
Mistral AI Mistral Small 4 20% off list price $0.15 our price $0.12 list price $0.60 our price $0.48 list price $0.015 our price $0.012 262K

Estimate your monthly API cost.

Estimate a chat or agent workload

Workload inputs

Estimated monthly API cost

$4.56 / month

Input
$1.68
Output
$2.88
Total token volume
10M

Planning estimate only. Actual charges follow metered usage and exclude tax and separately agreed services.

Product Plan

A unit is one model’s base allocation of output throughput. Choose how many units you keep all month, and add more for busy periods in 2-hour blocks.

Monthly base

$1,200 per unit, per model, per month

Reserved capacity on one model, available at every hour of the month.

Example: two units of GLM-5.3 are $2,400 a month.

Auto scaling

$5 per extra unit, per 2-hour period

Extra units you add on top of your base during a traffic peak, billed only for the 2-hour periods they run.

Example: one extra unit for a 6-hour peak is $15.

Estimate your monthly Product Plan cost.

One model, one peak event

Plan inputs

Estimated monthly cost

$1,205 / month

Monthly base 1 × $1,200
$1,200
This peak 1 extra unit × 1 period × $5
$5

Planning estimate, before tax. Extra peak periods in the month are charged separately. For an estimate from your own traffic, use the Product Plan capacity calculator.

What one busy day looks like.

Your base runs all month. Each peak block is one 2-hour period of auto scaling on top of it.

Monthly baseAuto scaling

Auto scaling requires your authorization and available capacity. Billing is monthly with a one-month minimum; adding units mid-month charges only the prorated difference for the rest of the month.

Product Plan base and auto scaling charges
Charge Price What it covers
Kimi K3 allocation $1,200 per unit, per model, per month 40 aggregate output tokens per second.
GLM-5.3 allocation $1,200 per unit, per model, per month 100 aggregate output tokens per second.
DeepSeek V4.1 Flash allocation $1,200 per unit, per model, per month 300 aggregate output tokens per second.
Three-model bundle $3,000 per month One allocation each of Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Save $600 per month compared with three individual plans.
Free extra capacity No extra charge Capacity above your base only while the shared pool has room. It is not reserved.
Auto scaling $5 per extra unit, per two-hour period Auto scaling adds capacity for traffic peaks after your authorization, subject to availability and plan terms.

Coding Plan

All plans include Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Compare allowance timing and busy-period priority as well as price.

$50 per developer, per month

  • Smaller allowance
  • 5-hour + weekly limits
  • Peak-time restrictions
  • Lowest priority
Get $50 plan

$100 per developer, per month

  • Larger allowance
  • Weekly limit only, no 5-hour limit
  • Peak-time restrictions
  • Priority above $50, below $200
Get $100 plan

Allowance amounts, supported tools, renewal dates, and cancellation terms are confirmed before activation. Peak-time restrictions apply to every tier, including $200: requests may slow, queue, or time out. Priority is $200 → $100 → $50; it is not a latency or availability guarantee.

Need help comparing the options?

Tell us whether you are building an application or choosing a coding plan.

API or Product Plan?

Estimate API cost from input and output tokens. For a Product Plan, choose a model allocation and account for auto scaling charges. Compare service conditions as well as the total. How to buy AI inference for a product on a monthly plan walks through the comparison.

What does Coding Plan include?

A monthly plan for one developer with Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Allowance timing and busy-period priority differ by tier. Check access and setup

Does an account activate a plan?

Account creation, product access, and plan activation are separate steps. Check access and plan conditions before activation.