Resources for choosing and using AI models

Practical tools and clear assumptions for your next AI workload.

Updated

Groq

$15.75

Fireworks

$13.95

Together

$18.00

Standard Thinking

$4.77

1,000 calls a day · 30 days · 50% cached input

Tool ·

What will GPT-OSS 120B cost?

Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.

27terms defined

Capacity units
11
Plan mechanics
8
Coding windows
5

A fourth group, reading a rate, adds three more

Guide ·

Capacity and plan terms, defined

A working glossary for capacity plans and coding plans: 27 terms, each quoted from the vendor's own page or from ours, with the ones we have not published marked.

6units for one allowance

Formula public
1 of 6
No figure
3 of 6
Open signup
5 of 6

Six vendors' published units, checked September 29, 2026

Comparison ·

How coding plans count usage: credits, requests, dollars, and windows

A coding plan's price is comparable; its allowance is not. What one unit counts at each vendor, which window refills it, and what the pages leave unsettled.

3ways to label text, compared

Tuned encoder
94.6 F1
LLM prompt
91.4 F1
Decision, 3B
84.5%

AG News; published studies and Standard One model cards

Comparison ·

Fine-tuned classifier or LLM: accuracy, speed and cost compared

With thousands of examples and a fixed label set, a fine-tuned encoder still wins on accuracy. Without them, or when labels change or need a probability, a decision model starts the same day at a similar cost per request.

2.6%worst case when 3 in 300 are wrong

Observed
1.0%
8B ECE
0.184
3B ECE
0.212

95% Clopper–Pearson bound; ECE from the Standard One model cards

Guide ·

How to set a confidence threshold for automated LLM decisions

A threshold decides which answers your code acts on alone. Set it on the chosen answer's probability, from a labelled sample of your own traffic, and measure again when the traffic changes.

9offerings compared

Price public
5 of 9
No minimum term
4 of 9
Unit is a GPU
3 of 9

Vendor terms read on September 16, 2026

Comparison ·

How LLM providers sell reserved capacity

Nine ways to buy inference capacity before you use it, put in one set of columns, with every term read from the vendor's own page.

46%utilization where API cost equals one unit

Model
GLM-5.3
Unit tok/s
100
API total
$2,592.00

Calculation example, not a measurement

Guide ·

How to buy AI inference for a product on a monthly plan

The four ways a product can pay for inference, what each one actually promises, and a worked break-even against our published rates.

900answers per hour

Throughput
100 tokens / s
Answer length
400 tokens
Utilization
100%

Calculation example, not a measurement

Guide ·

What “aggregate output tokens per second” means for your product

A throughput unit is a rate shared by every request at once. Read it as answers per hour, and size peaks in 2-hour blocks.

$3,000a month for the example product

Pay per token
$4,500
Base
2 GLM-5.3 units
Auto scaling
60 blocks

Calculation example, not a measurement

Tool ·

How many Product Plan units does your traffic need, and what would pay per token cost?

Size a monthly plan from your busiest ordinary hour, price the peaks in blocks, and read the month against pay per token.

19.6Mtokens a month

Input
18.8M
Output
850K
Model calls
5K

Simple tool loop · 1,000 runs a month

Tool ·

Plan tokens for chats and agents

Account for repeated conversation history and tool calls before choosing a rate card.