Resources for choosing and using AI models
Practical tools and clear assumptions for your next AI workload.
1,000 calls a day · 30 days · 50% cached input
What will GPT-OSS 120B cost?
Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.
27terms defined
- Capacity units
- 11
- Plan mechanics
- 8
- Coding windows
- 5
A fourth group, reading a rate, adds three more
Capacity and plan terms, defined
A working glossary for capacity plans and coding plans: 27 terms, each quoted from the vendor's own page or from ours, with the ones we have not published marked.
6units for one allowance
- Formula public
- 1 of 6
- No figure
- 3 of 6
- Open signup
- 5 of 6
Six vendors' published units, checked September 29, 2026
How coding plans count usage: credits, requests, dollars, and windows
A coding plan's price is comparable; its allowance is not. What one unit counts at each vendor, which window refills it, and what the pages leave unsettled.
3ways to label text, compared
- Tuned encoder
- 94.6 F1
- LLM prompt
- 91.4 F1
- Decision, 3B
- 84.5%
AG News; published studies and Standard One model cards
Fine-tuned classifier or LLM: accuracy, speed and cost compared
With thousands of examples and a fixed label set, a fine-tuned encoder still wins on accuracy. Without them, or when labels change or need a probability, a decision model starts the same day at a similar cost per request.
2.6%worst case when 3 in 300 are wrong
- Observed
- 1.0%
- 8B ECE
- 0.184
- 3B ECE
- 0.212
95% Clopper–Pearson bound; ECE from the Standard One model cards
How to set a confidence threshold for automated LLM decisions
A threshold decides which answers your code acts on alone. Set it on the chosen answer's probability, from a labelled sample of your own traffic, and measure again when the traffic changes.
9offerings compared
- Price public
- 5 of 9
- No minimum term
- 4 of 9
- Unit is a GPU
- 3 of 9
Vendor terms read on September 16, 2026
How LLM providers sell reserved capacity
Nine ways to buy inference capacity before you use it, put in one set of columns, with every term read from the vendor's own page.
46%utilization where API cost equals one unit
- Model
- GLM-5.3
- Unit tok/s
- 100
- API total
- $2,592.00
Calculation example, not a measurement
How to buy AI inference for a product on a monthly plan
The four ways a product can pay for inference, what each one actually promises, and a worked break-even against our published rates.
900answers per hour
- Throughput
- 100 tokens / s
- Answer length
- 400 tokens
- Utilization
- 100%
Calculation example, not a measurement
What “aggregate output tokens per second” means for your product
A throughput unit is a rate shared by every request at once. Read it as answers per hour, and size peaks in 2-hour blocks.
$3,000a month for the example product
- Pay per token
- $4,500
- Base
- 2 GLM-5.3 units
- Auto scaling
- 60 blocks
Calculation example, not a measurement
How many Product Plan units does your traffic need, and what would pay per token cost?
Size a monthly plan from your busiest ordinary hour, price the peaks in blocks, and read the month against pay per token.
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.