Skip to content

Estimate your monthly API cost.

Choose a request, chat, or agent workload. Adjust the starting values to match your product. This estimates API usage, not subscription capacity.

This tool estimates token-billed API workloads. It does not calculate Coding Plan subscription prices or Product Plan allocations for your product.

What does one workload look like?

Agent workflow structure

Choose a starting workflow, then adjust its model calls, tool results, and context assumptions. A compact loop that alternates model actions and tool observations.

History sent again is the share of earlier content included in later model calls. It does not change how Standard Thinking stores your data. A history cap of 0 means no cap in this estimate.

Model and cost assumptions

The multiplier repeats the entire scenario, including model and tool calls. Add model-only evaluation calls to the workflow assumptions instead. Rates are editable assumptions; view API rates before using the estimate for budgeting.

Three choices shape every estimate.

Choose a workload, then replace the starting assumptions with values from your application. Agent presets describe example workflows, not product benchmarks.

What does one workload look like?

Requests counted separately · Chat history · Tools and model calls

API · Chat · Agent

Workflow preset

A compact loop that alternates model actions and tool observations.

Simple tool loop

Model and cost assumptions

The workload multiplier scales the complete scenario, including model and tool calls, to represent retries or repeated runs.

Context limit · rates

Count what the model processes on each call.

Include instructions, earlier messages, plans, and tool results each time they are sent to the model. Count parallel workers as separate calls.

Include context sent again.

A run with five model calls counts its repeated prefix five times, so input grows with the number of calls.

Count each model call.

Planning, worker tasks, summaries, and retries may add calls to the workflow.

Check the largest context.

Monthly tokens help estimate cost. The largest request context helps assess model fit.

Use rates for your model.

Replace the example rates and context limit with the values that apply to your selected model.

Use the estimate to compare API rates. For a Product Plan, share your peak traffic and request-size assumptions with us.

View all resources