Docs · Standard One
How Standard One makes decisions
What its two sizes read and return, how to call them, and how we measured them.
DDana Whitfield
I was charged twice for order #48213. Please refund the extra charge…
Receipt photo1 image attached
+ 3 questions
routechoicerefundnoulurgencyscore
One 3B
standard-one-3b
3 answers
routeBilling.88refundYes.94urgencyHigh.58
What Standard One is
Standard One is Standard Thinking's decision model, in two sizes: Standard One 3B and Standard One 8B. Give it a state — a short piece of text, an image, or both — and one or more questions, each with the options you define, and it answers every question with one of those options and a probability. No paragraph to parse, no retry loop for a reply that came back in the wrong shape.
Both sizes take the same request, return the same response format, and read text and images. One 3B is the faster of the two and One 8B the more accurate; both are open weights, and both are served through the API. Choose a size compares them.
It's built for the narrow, closed-ended decisions a product makes over and over: which team should own a ticket, whether a photo is safe to publish, whether a message is a buying signal. Where teams use it has more examples. The model IDs are standard-one-3b and standard-one-8b.
This page is the technical reference. To see the models in use first, watch One 3B decide a support queue or read three worked examples on the Standard One page.
Choose a size
The choice comes down to speed, accuracy, and price:
| Model | Accuracy | Median time | Input / 1M tokens | Model ID | Weights |
|---|---|---|---|---|---|
| One 3B | 90.2% | 22.6 ms | $0.019 text · $0.036 image | standard-one-3b | Hugging Face |
| One 8B | 90.8% | 25.8 ms | $0.049 text · $0.084 image | standard-one-8b | Hugging Face |
Both read text and images and return the same answer fields. Accuracy is on a set of 600 realistic decisions; median time is the p50 latency on one NVIDIA H200. Both come from the evaluation below. Output tokens aren't billed.
The two sizes score about the same on realistic decisions. One 8B pulls ahead on hard cases: 54.95% against 47.75% on the JevBench public hard tier, and 76.6% against 69.9% averaged over ten suites. Start with One 8B for the decisions where accuracy matters most, and measure it on a sample of your own cases. Paths that run at high volume or on a tight latency budget can move to One 3B wherever accuracy holds on that same sample. On either size, keep a threshold, so a case the model isn't sure about goes to a person — see Probabilities and thresholds.
To run a model yourself, download its weights from Hugging Face: StandardThinking/StandardOne-3B or StandardThinking/StandardOne-8B, both under the Apache 2.0 license. The StandardThinking organization page also carries FP8, GGUF, and LoRA variants of both sizes.
What you send
state is the situation the model looks at: a string, or a JSON object or array, which it reads as JSON text. Pictures go beside it in images, a list of data URLs or HTTP(S) URLs. The request is the same on One 3B and One 8B.
| Form | Example |
|---|---|
| Text | "state": "Do you support SSO on the Team plan?" |
| Structured | "state": { "plan": "Team", "request": "SSO" } |
| Text and image | "state": "Receipt for an expense claim.", "images": ["data:image/jpeg;base64,…"] |
Every request needs a state, even when a picture carries the case; a short line such as "Receipt for an expense claim." is enough. Image limits are in Limits and pricing.
questions is a map of the questions you want answered about that state, in one call. Each one names its type (choice, noul, or score — the next section covers all three) and asks the question itself in instructions. criteria describes the options: for choice, a map from each option's name to one line about it; for score, a list of levels in order, the first being level 0; for noul, optionally, one line each for "true" and "false".
Three kinds of questions
Every question you ask is one of three types. All three come back in the same call, however many you send.
| Type | Ask it when | Returns | Example |
|---|---|---|---|
choice |
You want the model to pick one option from a list you define | choice, the most likely option, with probabilities for every option and a confidence |
"Who should pick this up?" → billing, technical, shipping, sales |
noul |
You want a yes-or-no judgment | noul, the probability that the answer is yes |
"Is the customer asking for a refund?" |
score |
You want a rating on an ordered scale you define | score, the expected level counting the first as 0, with probabilities for every level, a legend and a confidence |
"How urgent is it?" → low, medium, high |
We call a yes-or-no question noul in the API. It works like the other types with a fixed pair of outcomes instead of a list you supply: it takes instructions, can describe "true" and "false" in criteria, and returns the probability of yes as noul.
What comes back
One worked example, used the same way through the rest of this page: a support message about a duplicate charge, with three questions attached to it.
{
"model": "standard-one-3b",
"state": "I was charged twice for order #48213. Please refund the extra charge today.",
"questions": {
"route": {
"type": "choice",
"instructions": "Who should pick this up?",
"criteria": {
"billing": "Payments, charges, and refunds.",
"technical": "A bug or a broken feature.",
"shipping": "A delivery or fulfillment problem.",
"sales": "A new purchase or an upgrade."
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is it?",
"criteria": [
"Low: can wait a few days.",
"Medium: should be handled today.",
"High: needs attention right now."
]
}
}
}
One 3B answers all three questions in one response. Every answer carries its type back with its probabilities:
{
"model": "standard-one-3b",
"answers": {
"route": {
"type": "choice",
"probabilities": {
"billing": 0.88,
"technical": 0.07,
"shipping": 0.03,
"sales": 0.02
},
"confidence": 0.65,
"choice": "billing"
},
"refund": {
"type": "noul",
"noul": 0.94
},
"urgency": {
"type": "score",
"probabilities": {
"0": 0.03,
"1": 0.39,
"2": 0.58
},
"confidence": 0.28,
"score": 1.55,
"legend": {
"0": "Low: can wait a few days.",
"1": "Medium: should be handled today.",
"2": "High: needs attention right now."
}
}
},
"usage": {
"input_tokens": 214,
"output_tokens": 0
},
"metadata": {
"confidence_method": "1 - normalized_entropy",
"temperature": 0.9,
"temperature_by_type": { "choice": 0.9, "noul": 1.15, "score": 1.2 },
"evaluations": 3
}
}
| Type | Answer field | Probability field |
|---|---|---|
choice | choice — the most likely option | probabilities — one number per option |
noul | — | noul — the probability of yes |
score | score — the expected level, counting the first as 0; legend maps each level back to your wording | probabilities — one number per level, keyed "0", "1", … |
confidence is one minus the normalized entropy of the probabilities: 1 when all the weight sits on one option, 0 when it is spread evenly. It describes spread, not correctness, so set your threshold on the probabilities. metadata reports the temperatures applied and how many evaluations the call ran; options.temperature in a request overrides the default.
usage.input_tokens counts what you sent. Every answer is a short, fixed-shape value rather than generated text, so there is nothing to meter on the way out. How image inputs are counted is still being confirmed.
Several questions in one call
questions is a map: you choose the keys — route, refund, urgency, or anything else that reads well in your own code — and describe one question under each. The example above asks three questions about the same state and gets all three answers back together in one response, each under the same key it was asked with.
Because every answer is a short, fixed-shape value rather than a generated paragraph, asking three questions in one call costs less than three separate free-text generations would. Exactly how latency scales with each extra question is still being measured — see Limits and pricing.
Probabilities and thresholds
Every answer carries a probability: the model's estimate, given what's in state, that this option is the right one — not a measure of how clearly the question was asked, and not a guess dressed up as a number.
A calibrated model's probabilities mean what they say. Pull out every answer where it said "90% confident" — a well-calibrated model is right in about 90 of every 100 of them. A model that runs overconfident is right less often than its own numbers claim; one that runs underconfident is right more often than it admits.
Your code sets the threshold, not the model. Set the bar higher for decisions that are expensive to get wrong, and lower where a wrong guess costs little; anything below it is a signal to route the case to a person instead of forcing a choice. In the example above, a 60% bar lets your code act on route (88%) and refund (94%) directly, and sends urgency, which came back high at only 58%, to a person.
We report calibration two ways. ECE, expected calibration error: group every answer by the probability the model gave it, in bands (a 0.90–0.95 band, an 0.80–0.85 band, and so on); in each band, compare the model's average stated probability to how often it was actually right; ECE is the size of that gap, averaged across bands and weighted by how many answers fall in each one. Distance from the stated odds: on questions that state the true odds, the total-variation distance between the probabilities the model returns and those odds. Zero is perfect for both; lower is better — see them measured.
How it compares to a general-purpose model
A general-purpose model can do almost anything with language. Standard One does one thing: given a state and a small set of options, it picks. That trade shows up in four places.
| A general-purpose model | Standard One | |
|---|---|---|
| What comes back | A paragraph you still have to read or parse | One of the options you gave it |
| How sure it is | Usually described in words, if at all | A probability for every option |
| Format | You parse the reply and retry when it breaks | Always the shape you asked for |
| Speed and cost | Grows with how much it writes | One short answer per question, however many you ask |
A typed answer can't come back malformed — that's a guarantee about shape, not about truth. A miscalibrated or simply wrong answer can still arrive with a confident-looking probability attached. Treat the number as a calibrated estimate worth weighing against a threshold, not a promise that it's correct — see Probabilities and thresholds.
Where teams use it
Either size, with the questions you write. A few common shapes:
| Workflow | What it decides | Sample answer |
|---|---|---|
| Support triage | Which team should take this ticket | technical · 0.95 |
| Content moderation | Whether a photo can be published | publish · 0.98 |
| Lead scoring | How warm an inbound request is | hot · 0.63 |
| Expense approvals | Approve, reject, or ask for details | ask for details · 0.86 |
| Search relevance | Whether a passage answers the question | no · 0.88 |
| Agent routing | Which tool an agent should call next | check_calendar · 0.86 |
The sample answers show the shape of a reply; they aren't measured results. Two of these shapes are weak without fine-tuning: ten-way support triage and passage relevance for retrieval score 36% and 59% on One 8B. The support triage example on the Standard One page uses four options. For ten-way triage or retrieval relevance, fine-tune the open weights on your own labelled cases.
Call the API
Every request is a POST to /v1/systemone with your state and questions; the response carries one answer per question, in the shape described above. You'll need a Standard Thinking API key — see Create an API key if you don't have one yet.
curl https://api.standardthinking.ai/v1/systemone \
-H "Authorization: Bearer $STANDARD_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "standard-one-3b",
"state": "I was charged twice for order #48213. Please refund the extra charge today.",
"questions": {
"route": {
"type": "choice",
"instructions": "Who should pick this up?",
"criteria": {
"billing": "Payments, charges, and refunds.",
"technical": "A bug or a broken feature.",
"shipping": "A delivery or fulfillment problem.",
"sales": "A new purchase or an upgrade."
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is it?",
"criteria": [
"Low: can wait a few days.",
"Medium: should be handled today.",
"High: needs attention right now."
]
}
}
}'
import requests
response = requests.post(
"https://api.standardthinking.ai/v1/systemone",
headers={"Authorization": f"Bearer {STANDARD_API_KEY}"},
json={
"model": "standard-one-3b",
"state": (
"I was charged twice for order #48213. "
"Please refund the extra charge today."
),
"questions": {
"route": {
"type": "choice",
"instructions": "Who should pick this up?",
"criteria": {
"billing": "Payments, charges, and refunds.",
"technical": "A bug or a broken feature.",
"shipping": "A delivery or fulfillment problem.",
"sales": "A new purchase or an upgrade.",
},
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
"urgency": {
"type": "score",
"instructions": "How urgent is it?",
"criteria": [
"Low: can wait a few days.",
"Medium: should be handled today.",
"High: needs attention right now.",
],
},
},
},
)
answers = response.json()["answers"]
const response = await fetch(
"https://api.standardthinking.ai/v1/systemone",
{
method: "POST",
headers: {
"Authorization": `Bearer ${STANDARD_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "standard-one-3b",
state:
"I was charged twice for order #48213. " +
"Please refund the extra charge today.",
questions: {
route: {
type: "choice",
instructions: "Who should pick this up?",
criteria: {
billing: "Payments, charges, and refunds.",
technical: "A bug or a broken feature.",
shipping: "A delivery or fulfillment problem.",
sales: "A new purchase or an upgrade.",
},
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?",
},
urgency: {
type: "score",
instructions: "How urgent is it?",
criteria: [
"Low: can wait a few days.",
"Medium: should be handled today.",
"High: needs attention right now.",
],
},
},
}),
},
);
const { answers } = await response.json();
Set model to standard-one-3b or standard-one-8b; nothing else in the request changes.
Python and JavaScript SDKs, CSV batch processing, and a LangChain integration are planned.
Deployment and data
Both sizes are called the same way as any other Standard Thinking model, over HTTPS, from the regions we publish for the API. Both can also be run from their open weights, linked in Choose a size. A dedicated deployment — capacity reserved for one team rather than shared — is planned for teams that need one.
"Dedicated" can mean several different things — separate capacity, a private network path, or data that never leaves a chosen region. We'll say exactly which of these a dedicated deployment provides once it's confirmed, rather than imply all of them.
Evaluation
We measured both sizes in the release configuration: merged BF16 weights behind SGLang 0.5.20 and jev-adapter, native prompt wording, no system prompt, one option order, and the temperature fitted for each answer type. The wording and the temperatures were chosen on data outside JevBench. Jev 1.13 was measured on the same questions through its hosted endpoint, with its raw probabilities. These are our measurements, not official sealed-set JevBench scores.
These results compare models on one narrow kind of task — picking from a short list of options, under a strict response format — not on writing or conversation generally. A general-purpose model given room to answer in its own words, or wired up with its own constrained decoding, can close some of this gap on an equivalent narrow task. Read the comparison on that basis, not as a general capability ranking.
Accuracy on the served endpoint (higher is better)
| Suite (questions) | One 3B | One 8B | Jev 1.13 |
|---|---|---|---|
| JevBench public, easy (48) | 100.00% | 100.00% | 100.00% |
| JevBench public, standard (72) | 86.11% | 93.06% | 98.61% |
| JevBench public, hard (111) | 47.75% | 54.95% | 72.07% |
| Judge proxy: routing and answer adequacy (600) | 87.50% | 89.33% | 90.50% |
| Realistic transfer set (600) | 90.17% | 90.83% | 86.67% |
| Stated-distribution probability (1,036) | 78.38% | 81.18% | 72.97% |
| Hard proxy (600) | 44.17% | 53.00% | 54.83% |
Latency, p50 (ms, lower is better)
One 3B
22.6 ms
One 8B
25.8 ms
Probability quality (lower is better)
| Measure | One 3B | One 8B | Jev 1.13 |
|---|---|---|---|
| Distance from the stated odds (total variation) | 0.124 | 0.108 | 0.192 |
| ECE on JevBench public, hard (111) | 0.212 | 0.184 | 0.099 |
Public classification suites (400 cases each, higher is better)
| Suite | One 3B | One 8B |
|---|---|---|
| Topics, AG News | 84.5% | 84.2% |
| Typed decisions | 67.3% | 71.1% |
| Intent, MASSIVE, mean of nine languages | 83.4% | 86.1% |
| Email spam | 92.0% | 91.8% |
| Phishing | 88.8% | 85.2% |
- Versions: One 3B v2.1, released September 27, 2026; One 8B v2, released September 26, 2026. One 8B was measured from September 24 to 26, 2026.
- Every accuracy figure is for text decisions. Accuracy on image inputs hasn't been benchmarked yet.
- A sealed-set JevBench v1.4.1 run has been requested; its result isn't available yet.
- Full tables, per-language results, offline runs and the server code are on the model cards: StandardOne-3B and StandardOne-8B.
Limits and pricing
| Limit | Value |
|---|---|
| Questions per call | Up to 256 |
Options per choice | 2 to 26 |
Levels per score | 2 to 10 |
| Context | 8,192 tokens |
| Image format | Data URLs (data:image/…;base64,…) or HTTP(S) URLs, in images |
| Image size | To be published |
| Languages | English, Japanese, Chinese, Spanish, French, German, Portuguese and Russian, with a smaller share of Korean; accuracy varies by language |
| Request rate | See general rate limits |
| Price, One 3B | $0.019 per 1M text input tokens, $0.036 per 1M image input tokens |
| Price, One 8B | $0.049 per 1M text input tokens, $0.084 per 1M image input tokens |
| Output tokens | Not billed on either size |
Frequently asked questions
Does Standard One read images?
Yes, both sizes do. Send pictures in images beside a short state — see What you send. The published accuracy figures are for text; accuracy on images hasn't been benchmarked yet.
Can I run Standard One myself?
Yes. Both sizes are open weights, published on Hugging Face as StandardThinking/StandardOne-3B and StandardThinking/StandardOne-8B, and both are also available through the API — see Choose a size. The server code in the StandardOne-8B repository answers on the same path, /v1/systemone, with the same request and response.
What does noul mean?
It's our API's name for a yes-or-no question. It returns noul, the probability that the answer is yes. If the question needs it, say what yes and no mean in criteria under "true" and "false".
Can Standard One write free text?
No. Every answer is one of the types you asked for — a choice, a yes-or-no, or a score — never a paragraph. For open-ended writing or conversation, use a general-purpose model instead.
What happens when it isn't sure?
Your own code decides. Choose the probability above which it acts on its own; below that, route the case to a person instead of forcing a choice — see Probabilities and thresholds.
Is there a live demo?
Yes. The API is live: sign in and send it a request — see Call the API. To see it in use first, watch One 3B decide a support queue, or read three worked examples on the Standard One page. The messages and answers in both are illustrations; measured results are in Evaluation.
Are there SDKs?
Python and JavaScript SDKs, CSV batch processing, and a LangChain integration are planned; package names and availability are being confirmed.