How to set a confidence threshold for automated LLM decisions
A threshold decides which answers your code acts on alone. Set it on the chosen answer's probability, from a labelled sample of your own traffic, and measure again when the traffic changes.
A threshold decides which answers your code acts on alone and which it hands to a person. Set it on the probability of the answer the model chose, read the value off a labelled sample of your own traffic, and measure it again when that traffic changes. A model's published calibration tells you how far to trust its numbers before you have that sample; it does not replace it.
Compare the chosen answer's probability with the bar
Standard One returns a probability for every option you supply. For a choice question, the number to compare with your bar is the probability of the option it picked: in the documentation's example, 0.88 for billing. The same answer also carries confidence, one minus the normalized entropy of all the probabilities, which read 0.65 there because the other three options still held some weight. The first number estimates whether this answer is right; the second measures how spread out the model's belief is. Published calibration figures describe the first.
A yes-or-no (noul) question returns the probability of yes, so an answer of no carries one minus that. A score question returns the expected level and a probability for each level: threshold the probability of the level you act on, or the sum over a band of levels.
Check how far the numbers can be trusted
Calibration asks whether 0.9 means right nine times in ten. Expected calibration error (ECE) sorts answers into ten probability bands and averages the gap between stated probability and observed accuracy, weighted by how many answers fall in each band. On the public JevBench hard tier, where both sizes answer about half the questions correctly, ECE at the served temperatures is 0.184 for Standard One 8B and 0.212 for 3B. On questions that state their true odds, the mean total-variation distance between the returned probabilities and those odds is 0.108 and 0.124.
Both model cards attach the condition that matters most for thresholds: the temperatures that shape these probabilities were fitted on Standard Thinking's calibration suites, and on a materially different distribution of questions they should be refitted rather than assumed to transfer. Treat published figures as a starting estimate, not as your error rate.
Read the threshold off a labelled sample
- Label a sample of real requests, drawn the way production traffic arrives, for each question you want to automate.
- Run every case and keep two values: the probability of the chosen answer, and whether it was right.
- For each candidate threshold, count the cases at or above it, which your code would handle alone, and the errors among them.
- Pick the lowest threshold whose error rate meets your target, judged by an upper confidence bound rather than the raw rate. If 300 answers clear the bar and 3 are wrong, the observed error is 1.0%, but the one-sided 95% Clopper–Pearson bound is 2.6%. Badshah, Emami and Sajjad (2026) use the same bound to keep the error among accepted LLM verdicts below a chosen level.
- Send everything below the threshold to review, and add the reviewed cases to the sample.
import csv
# labelled.csv: the chosen answer's probability, and 1 if the answer was right
rows = [(float(r["probability"]), r["correct"] == "1") for r in csv.DictReader(open("labelled.csv"))]
for bar in (0.5, 0.6, 0.7, 0.8, 0.9, 0.95):
kept = [ok for p, ok in rows if p >= bar]
errors = kept.count(False)
print(f"{bar:.2f} handled alone {len(kept) / len(rows):.1%} errors {errors}/{len(kept)}")
No errors in the sample does not mean no errors in production: with none among 300 accepted answers, the same bound is still 1.0%.
Price both sides of the line
Every answer above the bar costs a model call and carries the chance of a wrong action; every answer below it costs a review. For each candidate threshold, the cost per case is the share handled alone times its error rate times the cost of a wrong action, plus the share reviewed times the cost of a review.
The model call is the smallest term. At the roughly 280 input tokens per decision used in the published latency test, a million decisions cost $5.32 on Standard One 3B and $13.72 on 8B at the current rates of $0.019 and $0.049 per million text input tokens, half the list price; output is free. The threshold is set by what a mistake and a review cost you, not by the model bill.
Keep measuring after launch
- Run the sample again when the inputs change: a new product line, a new language, reworded options.
- Watch the share sent to review. A jump means the traffic moved, even if the error rate above the bar still looks fine.
- Audit a small random slice of automated answers every week, so the error rate above the bar stays measured rather than assumed.
- Some tasks no threshold rescues. The 8B card reports ten-way support triage at 36% and RAG passage relevance at 59% without task training, and advises fine-tuning for those; when a fine-tuned classifier beats an LLM covers that choice.
Standard Thinking serves Standard One; its request shapes and measured results are on the model page and in the documentation.
Sources and verification
Sources used for this resource. Verification dates describe our source checks, not provider effective dates.
Standard Thinking documentation, Standard One answers and thresholds
The worked answer quoted in the first section: route returns probabilities billing 0.88, technical 0.07, shipping 0.03, sales 0.02 with "confidence": 0.65 and "confidence_method": "1 - normalized_entropy"; noul returns the probability of yes; score returns the expected level, a probability per level and a legend. The thresholds section states "Your code sets the threshold, not the model" and defines ECE by probability bands and the distance from stated odds by total variation. The answer shapes follow the Hugging Face adapter at POST /v1/systemone, documented since September 29, 2026.
Standard One 8B model card (v2)
Served endpoint, per-answer-type temperatures (choice 0.85, noul 0.85, score 0.70): "Hard-tier ECE at the served temperatures is 0.184 against Jev 1.13's 0.099 raw, and mean TV to the stated distributions is 0.108 at served T"; public hard tier 54.95%. The refit condition: "if you apply this model to a materially different question distribution, re-fitting that temperature is advisable rather than assuming these values transfer." The weak workflows: "Ten-way support triage (36 %) and RAG passage relevance (59 %) are weak zero-shot; fine-tune for those." Latency: raw serial p50 on one H200 with SGLang 0.5.20 for "a 242-decision profile averaging about 280 input tokens per decision".
Standard One 3B model card (v2.1)
"At the fitted temperatures, public-hard 10-bin ECE is 0.212 and mean total-variation distance on the stated-distribution validation set is 0.124." Served public hard tier 47.75%. "Temperatures were fitted separately for each wording and answer type on three calibration suites only", and the limitation "Refit temperatures when using a materially different distribution."
Standard One benchmark report, calibration and latency
Used for definitions only: the ECE reported is "Hard-tier ECE (10-bin, top label)", that is, the probability of the chosen answer, and the latency profile is 242 decisions at about 280 input tokens each, run serially on one H200. The report's own figures are from v1.1 and are not quoted; the v2 and v2.1 figures above come from the model cards.
Badshah, Emami and Sajjad, Judge, Retrieve, or Abstain (arXiv 2608.17994)
Submitted 18 August 2026. The abstract describes a framework that "calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level" using "finite-sample Clopper–Pearson intervals". The 2.6% and 1.0% figures on this page are our exact one-sided 95% Clopper–Pearson upper bounds for 3 errors in 300 and 0 errors in 300, not figures from the paper.
Standard Thinking pricing, Standard One
The rates used for the per-decision cost: text input at $0.019 per 1M tokens for One 3B and $0.049 for One 8B, each shown as 50% off list prices of $0.038 and $0.098, and the sentence "Choose a size and pay per input token. Each call returns one short answer, so output is free."
Continue exploring
3ways to label text, compared
- Tuned encoder
- 94.6 F1
- LLM prompt
- 91.4 F1
- Decision, 3B
- 84.5%
AG News; published studies and Standard One model cards
Fine-tuned classifier or LLM: accuracy, speed and cost compared
With thousands of examples and a fixed label set, a fine-tuned encoder still wins on accuracy. Without them, or when labels change or need a probability, a decision model starts the same day at a similar cost per request.
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.
1,000 calls a day · 30 days · 50% cached input
What will GPT-OSS 120B cost?
Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.