What will GPT-OSS 120B cost?

Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.

By Standard Thinking · Updated

Prices checked . Groq: $0.15 input, $0.60 output, $0.075 cached input per million tokens; Fireworks: $0.15 input, $0.60 output, $0.015 cached input per million tokens; Together: $0.15 input, $0.60 output per million tokens, cached input not listed; Standard Thinking: $0.037 input, $0.17 output per million tokens, cached input not listed.

Estimate token cost for your workload

USD per 1M tokens · No API requests are made

Your workload

An assumed share of tokens, not a predicted request hit rate. Understand caching.

Estimated token cost

30,000 calls · 60,000,000 input · 15,000,000 output tokens

Groq · Standard API

Uncached input$9.00
Cached input$0.00
Output$9.00
Total$18.00

Total for 30 days. Uses the cached-input share above.

Fireworks · Standard serverless

Uncached input$9.00
Cached input$0.00
Output$9.00
Total$18.00

Total for 30 days. Uses the cached-input share above.

Together · Serverless

Uncached input$9.00
Cached inputNot modeled
Output$9.00
Total$18.00

Total for 30 days. Cache price not listed; all input uses the base rate.

Standard Thinking · Standard API

Uncached input$2.22
Cached inputNot modeled
Output$2.55
Total$4.77

Total for 30 days. Cache price not listed; all input uses the base rate.

Differences reflect the listed rates and your cache assumption. An unverified cache price is estimated at the normal input rate.

Prices checked . This is a dated rate snapshot; confirm current terms before budgeting.

Published token rates

Input, cached-input, and output rates used in the calculator.
Provider / tierInputCached inputOutputSource
GroqStandard API$0.15$0.075$0.60Rate source
FireworksStandard serverless$0.15$0.015$0.60Rate source
TogetherServerless$0.15Not listed$0.60Rate source
Standard ThinkingStandard API$0.037Not listed$0.17Rate source

USD per 1 million tokens · GPT-OSS 120B · standard tiers

Groq

The discounted rate applies to successfully cached input tokens. A matching prefix does not guarantee a hit.

openai/gpt-oss-120b

Fireworks

Uses the standard serverless cached-input rate. Priority, batch, and location-specific premiums are outside this comparison.

accounts/fireworks/models/gpt-oss-120b

Together

The source does not list a cached-input price for this model. This estimate uses the standard input rate for all input tokens; it does not establish that caching is unsupported.

openai/gpt-oss-120b

Standard Thinking

Our pricing page lists no cache-read rate for this model, so the estimate charges every input token at the input rate.

gpt-oss-120b

How to read the comparison

Start with the number of API calls, rather than the number of people using your app. One conversation or agent task can make several calls. Include any earlier messages or tool results sent again in the input-token estimate. For longer workflows, use the chat and agent token estimator first.

The table compares published token charges for the same named model on four providers, one of them our own Standard API. Matching model names do not establish identical deployments, output lengths, quality, latency, or service limits. This is a cost estimate, not a performance benchmark or a provider ranking.

What the cache assumption means

“Cached input” is the percentage of total input tokens you assume will actually be billed at a cached-input rate. It is not the percentage of requests that share a prompt, and the calculator does not predict whether a cache hit will occur.

The same assumption is applied to providers with a verified cached-input rate so that you can inspect the pricing difference. Actual cached shares can differ between providers. Check each provider’s matching rules and eligibility before using this scenario as a budget. When a cached rate is unverified, the estimate uses the normal input rate and labels that limitation.

The calculation

For each provider:

calls = requests per day × days
input = calls × average input tokens
output = calls × average billed output tokens
cached input = input × assumed cached-input share

cost = (uncached input × input rate
      + cached input × cached-input rate
      + output × output rate) / 1,000,000

Rates are in USD per million tokens. Calculation uses unrounded values; only the displayed currency is rounded. Reset restores the example workload. Its values are planning assumptions, not observed customer usage.

What to include in your token counts

Use the provider’s billed output usage, including reasoning tokens when they are charged as output. Include repeated calls and retries in request volume. The estimator cannot infer these from the text visible to an end user.

This comparison covers input, cached-input, and output token charges. It excludes taxes, credits, negotiated discounts, tool execution, storage, networking, and any separately charged cache-write or retention fees. Batch, priority, reserved capacity, and subscriptions need their own rate cards.

From an estimate to a budget

Run a representative workload with the provider you are considering, then replace the example values with measured token usage and the cached share in your billing records. Check success rate and response time alongside cost. A low price per token does not by itself tell you the cost of completing a task successfully.

The Standard API serves gpt-oss-120b at the rate in the table. Log in to run the same workload against it.

Sources and verification

Sources used for this resource. Verification dates describe our source checks, not provider effective dates.

Groq: GPT-OSS 120B pricing

The model page lists $0.15 input, $0.075 cached input, and $0.60 output per million tokens.

Checked

Groq: prompt caching conditions

GPT-OSS 120B is listed as supported. The discount requires a successful prefix cache hit; hits are not guaranteed.

Checked

Fireworks: serverless pricing

The standard serverless GPT OSS 120B row is $0.15 input, $0.015 cached input, and $0.60 output per million tokens. Priority pricing is a separate column and is not used.

Checked

Together: serverless model prices

The GPT-OSS 120B row lists $0.15 input and $0.60 output per million tokens; cached input is a dash. No cache discount is inferred.

Checked

Standard Thinking pricing, Standard API

gpt-oss-120b at $0.037 input and $0.17 output per 1M tokens; the cache-read column shows no rate for this model.

Checked

View all resources