Product Plan

Run your product’s AI on a monthly plan

Reserved base capacity for your product, agent, or service. No monthly token cap. Add capacity when traffic grows.

Choose one model or all three.

Each model has a guaranteed base output rate shared across your product’s concurrent requests.

Kimi K3

40 tokens/s

GLM-5.3

100 tokens/s

DeepSeek V4.1 Flash

300 tokens/s

Your order defines the model, guaranteed base allocation, measurement conditions, and any service commitments. The throughput figures do not describe a single request’s speed or a monthly token allowance. Billing is monthly with a one-month minimum; adding units mid-month charges only the prorated difference for the rest of the month.

Keep your base.
Add capacity when you need it.

Your base stays available during busy periods under the conditions in your order. Free extra capacity is available while the shared pool has room. Auto scaling requires your authorization and available capacity.

Product Plan base and auto scaling prices
Plan component Price What it covers
One model allocation $1,200 per unit, per model, per month Kimi K3, GLM-5.3, or DeepSeek V4.1 Flash. Includes that model’s base allocation. Add units of the same model to scale that base in $1,200 steps.
Three-model bundle $3,000 per month One allocation each of Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Save $600 per month compared with three individual plans.
Free extra capacity No extra charge Capacity above your base only while the shared pool has room. It is not reserved.
Auto scaling $5 per extra unit, per two-hour period Auto scaling adds capacity for traffic peaks after your authorization, subject to availability and plan terms.

One unit is that model’s base allocation shown above.

Product Plan is reserved capacity, sold as a monthly plan.

Inference is usually sold per token or by the hour, and a Product Plan unit is that capacity at a published monthly price.

Ways to buy inference capacity
How you buy What you pay for Where it fits
Pay per token (Standard API) Input and output tokens as you use them Variable or early traffic
Dedicated GPUs by the hour A GPU, whether or not it is busy Teams that run their own serving stack
Provisioned throughput by the hour A throughput unit metered per hour, usually on a reservation Enterprise contracts sized in tokens per minute
Product Plan One model’s output throughput for a whole month at a published price, plus $5 per extra unit, per two-hour period A product whose AI runs all month with peaks

Connect your product.

Create an account to get started. The workload, plan, and integration details below are confirmed before a Product Plan begins.

Share the workload

Include normal and peak traffic, typical input and output lengths, and required response behavior.

Review the plan

Review the model, base capacity, auto scaling charges, data terms, renewal, and cancellation.

Prepare the integration

Agree on authentication, environments, monitoring, and the activation process.

Store your product’s data with us. Keep control of it.

Product Plan can store your product’s data on Standard Thinking. Storage stays off until you turn it on; you control access to it and use it to improve your service.

Your application sends requests to Standard Thinking inference, with retention governed by the applicable service terms. Data you designate is stored on the data platform for your use only, and storage is optional. Standard Thinking processes it to provide your requested services, not for its own independent purposes. Your application sends requests to Standard Thinking inference, with retention governed by the applicable service terms. Data you designate is stored on the data platform for your use only, and storage is optional. Standard Thinking processes it to provide your requested services, not for its own independent purposes.

Your data is for your use only, and storage is optional.

Data control

The data platform is off until you turn it on. You control access to your data and how it is used.

Operational metadata

Content-free metadata is kept for billing, reliability, and service operations.

Model training

We do not train public or shared models on your inputs or outputs without separate, explicit opt-in. For how other providers state retention and training, see what inference providers keep after a request.

Questions about Product Plans.

Choose by the workload and service conditions you need.

How is this different from API?

Both power AI features in your product. API bills input and output tokens. Product Plan includes monthly base capacity with no monthly token cap, free extra capacity when available, and optional auto scaling. Product Plan also includes a data platform where your product’s data is stored under your control. The Product Plan capacity calculator sizes a plan from your traffic and prices the same month per token. Compare API rates

Will each answer arrive faster?

Response time depends on the model, request size, concurrency, and service conditions. Review the defined base commitment and any latency commitment.

Who manages the end user?

Your application manages users, permissions, and the customer experience. The integration plan defines service responsibilities.

Can I resell the allocation?

A Product Plan covers the AI features inside your own product. Any other use, including resale, is permitted only where your order and the selected model’s license allow it.