Groq · Standard API
What will GPT-OSS 120B cost?
Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.
Prices checked . Groq: $0.15 input, $0.60 output, $0.075 cached input per million tokens; Fireworks: $0.15 input, $0.60 output, $0.015 cached input per million tokens; Together: $0.15 input, $0.60 output per million tokens, cached input not listed; Standard Thinking: $0.037 input, $0.17 output per million tokens, cached input not listed.
Estimate token cost for your workload
Fireworks · Standard serverless
Together · Serverless
Standard Thinking · Standard API
Prices checked . This is a dated rate snapshot; confirm current terms before budgeting.
Published token rates
| Provider / tier | Input | Cached input | Output | Source |
|---|---|---|---|---|
| GroqStandard API | $0.15 | $0.075 | $0.60 | Rate source |
| FireworksStandard serverless | $0.15 | $0.015 | $0.60 | Rate source |
| TogetherServerless | $0.15 | Not listed | $0.60 | Rate source |
| Standard ThinkingStandard API | $0.037 | Not listed | $0.17 | Rate source |
USD per 1 million tokens · GPT-OSS 120B · standard tiers
Groq
The discounted rate applies to successfully cached input tokens. A matching prefix does not guarantee a hit.
openai/gpt-oss-120b
Fireworks
Uses the standard serverless cached-input rate. Priority, batch, and location-specific premiums are outside this comparison.
accounts/fireworks/models/gpt-oss-120b
Together
The source does not list a cached-input price for this model. This estimate uses the standard input rate for all input tokens; it does not establish that caching is unsupported.
openai/gpt-oss-120b
Standard Thinking
Our pricing page lists no cache-read rate for this model, so the estimate charges every input token at the input rate.
gpt-oss-120b
How to read the comparison
Start with the number of API calls, rather than the number of people using your app. One conversation or agent task can make several calls. Include any earlier messages or tool results sent again in the input-token estimate. For longer workflows, use the chat and agent token estimator first.
The table compares published token charges for the same named model on four providers, one of them our own Standard API. Matching model names do not establish identical deployments, output lengths, quality, latency, or service limits. This is a cost estimate, not a performance benchmark or a provider ranking.
What the cache assumption means
“Cached input” is the percentage of total input tokens you assume will actually be billed at a cached-input rate. It is not the percentage of requests that share a prompt, and the calculator does not predict whether a cache hit will occur.
The same assumption is applied to providers with a verified cached-input rate so that you can inspect the pricing difference. Actual cached shares can differ between providers. Check each provider’s matching rules and eligibility before using this scenario as a budget. When a cached rate is unverified, the estimate uses the normal input rate and labels that limitation.
The calculation
For each provider:
calls = requests per day × days
input = calls × average input tokens
output = calls × average billed output tokens
cached input = input × assumed cached-input share
cost = (uncached input × input rate
+ cached input × cached-input rate
+ output × output rate) / 1,000,000
Rates are in USD per million tokens. Calculation uses unrounded values; only the displayed currency is rounded. Reset restores the example workload. Its values are planning assumptions, not observed customer usage.
What to include in your token counts
Use the provider’s billed output usage, including reasoning tokens when they are charged as output. Include repeated calls and retries in request volume. The estimator cannot infer these from the text visible to an end user.
This comparison covers input, cached-input, and output token charges. It excludes taxes, credits, negotiated discounts, tool execution, storage, networking, and any separately charged cache-write or retention fees. Batch, priority, reserved capacity, and subscriptions need their own rate cards.
From an estimate to a budget
Run a representative workload with the provider you are considering, then replace the example values with measured token usage and the cached share in your billing records. Check success rate and response time alongside cost. A low price per token does not by itself tell you the cost of completing a task successfully.
The Standard API serves gpt-oss-120b at the rate in the table. Log in to run the same workload against it.
Sources and verification
Sources used for this resource. Verification dates describe our source checks, not provider effective dates.
The model page lists $0.15 input, $0.075 cached input, and $0.60 output per million tokens.
Groq: prompt caching conditions
GPT-OSS 120B is listed as supported. The discount requires a successful prefix cache hit; hits are not guaranteed.
The standard serverless GPT OSS 120B row is $0.15 input, $0.015 cached input, and $0.60 output per million tokens. Priority pricing is a separate column and is not used.
Together: serverless model prices
The GPT-OSS 120B row lists $0.15 input and $0.60 output per million tokens; cached input is a dash. No cache discount is inferred.
Standard Thinking pricing, Standard API
gpt-oss-120b at $0.037 input and $0.17 output per 1M tokens; the cache-read column shows no rate for this model.
Continue exploring
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.
27terms defined
- Capacity units
- 11
- Plan mechanics
- 8
- Coding windows
- 5
A fourth group, reading a rate, adds three more
Capacity and plan terms, defined
A working glossary for capacity plans and coding plans: 27 terms, each quoted from the vendor's own page or from ours, with the ones we have not published marked.
6units for one allowance
- Formula public
- 1 of 6
- No figure
- 3 of 6
- Open signup
- 5 of 6
Six vendors' published units, checked September 29, 2026
How coding plans count usage: credits, requests, dollars, and windows
A coding plan's price is comparable; its allowance is not. What one unit counts at each vendor, which window refills it, and what the pages leave unsettled.