Skip to content

Docs · Standard One

How Standard One makes decisions

What its two sizes read and return, how to call them, and how we measured them.

Technical write-up · · Updated · 12 min read

State and questions in; one typed answer per question, with a probability, out.

What Standard One is

Standard One is Standard Thinking's decision model, in two sizes: Standard One 3B and Standard One 8B. Give it a state — a short piece of text, an image, or both — and one or more questions, each with the options you define, and it answers every question with one of those options and a probability. No paragraph to parse, no retry loop for a reply that came back in the wrong shape.

Both sizes take the same request, return the same response format, and read text and images. One 3B is the faster of the two and One 8B the more accurate; both are open weights, and both are served through the API. Choose a size compares them.

It's built for the narrow, closed-ended decisions a product makes over and over: which team should own a ticket, whether a photo is safe to publish, whether a message is a buying signal. Where teams use it has more examples. The model IDs are standard-one-3b and standard-one-8b.

This page is the technical reference. To see the models in use first, watch One 3B decide a support queue or read three worked examples on the Standard One page.

Choose a size

The choice comes down to speed, accuracy, and price:

ModelAccuracyMedian timeInput / 1M tokensModel IDWeights
One 3B90.2%22.6 ms$0.019 text · $0.036 imagestandard-one-3bHugging Face
One 8B90.8%25.8 ms$0.049 text · $0.084 imagestandard-one-8bHugging Face

Both read text and images and return the same answer fields. Accuracy is on a set of 600 realistic decisions; median time is the p50 latency on one NVIDIA H200. Both come from the evaluation below. Output tokens aren't billed.

The two sizes score about the same on realistic decisions. One 8B pulls ahead on hard cases: 54.95% against 47.75% on the JevBench public hard tier, and 76.6% against 69.9% averaged over ten suites. Start with One 8B for the decisions where accuracy matters most, and measure it on a sample of your own cases. Paths that run at high volume or on a tight latency budget can move to One 3B wherever accuracy holds on that same sample. On either size, keep a threshold, so a case the model isn't sure about goes to a person — see Probabilities and thresholds.

To run a model yourself, download its weights from Hugging Face: StandardThinking/StandardOne-3B or StandardThinking/StandardOne-8B, both under the Apache 2.0 license. The StandardThinking organization page also carries FP8, GGUF, and LoRA variants of both sizes.

What you send

state is the situation the model looks at: a string, or a JSON object or array, which it reads as JSON text. Pictures go beside it in images, a list of data URLs or HTTP(S) URLs. The request is the same on One 3B and One 8B.

FormExample
Text"state": "Do you support SSO on the Team plan?"
Structured"state": { "plan": "Team", "request": "SSO" }
Text and image"state": "Receipt for an expense claim.", "images": ["data:image/jpeg;base64,…"]

Every request needs a state, even when a picture carries the case; a short line such as "Receipt for an expense claim." is enough. Image limits are in Limits and pricing.

questions is a map of the questions you want answered about that state, in one call. Each one names its type (choice, noul, or score — the next section covers all three) and asks the question itself in instructions. criteria describes the options: for choice, a map from each option's name to one line about it; for score, a list of levels in order, the first being level 0; for noul, optionally, one line each for "true" and "false".

Three kinds of questions

Every question you ask is one of three types. All three come back in the same call, however many you send.

TypeAsk it whenReturnsExample
choice You want the model to pick one option from a list you define choice, the most likely option, with probabilities for every option and a confidence "Who should pick this up?" → billing, technical, shipping, sales
noul You want a yes-or-no judgment noul, the probability that the answer is yes "Is the customer asking for a refund?"
score You want a rating on an ordered scale you define score, the expected level counting the first as 0, with probabilities for every level, a legend and a confidence "How urgent is it?" → low, medium, high

We call a yes-or-no question noul in the API. It works like the other types with a fixed pair of outcomes instead of a list you supply: it takes instructions, can describe "true" and "false" in criteria, and returns the probability of yes as noul.

What comes back

One worked example, used the same way through the rest of this page: a support message about a duplicate charge, with three questions attached to it.

POST /v1/systemone
{
  "model": "standard-one-3b",
  "state": "I was charged twice for order #48213. Please refund the extra charge today.",
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Who should pick this up?",
      "criteria": {
        "billing": "Payments, charges, and refunds.",
        "technical": "A bug or a broken feature.",
        "shipping": "A delivery or fulfillment problem.",
        "sales": "A new purchase or an upgrade."
      }
    },
    "refund": {
      "type": "noul",
      "instructions": "Is the customer asking for a refund?"
    },
    "urgency": {
      "type": "score",
      "instructions": "How urgent is it?",
      "criteria": [
        "Low: can wait a few days.",
        "Medium: should be handled today.",
        "High: needs attention right now."
      ]
    }
  }
}

One 3B answers all three questions in one response. Every answer carries its type back with its probabilities:

200JSON
{
  "model": "standard-one-3b",
  "answers": {
    "route": {
      "type": "choice",
      "probabilities": {
        "billing": 0.88,
        "technical": 0.07,
        "shipping": 0.03,
        "sales": 0.02
      },
      "confidence": 0.65,
      "choice": "billing"
    },
    "refund": {
      "type": "noul",
      "noul": 0.94
    },
    "urgency": {
      "type": "score",
      "probabilities": {
        "0": 0.03,
        "1": 0.39,
        "2": 0.58
      },
      "confidence": 0.28,
      "score": 1.55,
      "legend": {
        "0": "Low: can wait a few days.",
        "1": "Medium: should be handled today.",
        "2": "High: needs attention right now."
      }
    }
  },
  "usage": {
    "input_tokens": 214,
    "output_tokens": 0
  },
  "metadata": {
    "confidence_method": "1 - normalized_entropy",
    "temperature": 0.9,
    "temperature_by_type": { "choice": 0.9, "noul": 1.15, "score": 1.2 },
    "evaluations": 3
  }
}
TypeAnswer fieldProbability field
choicechoice — the most likely optionprobabilities — one number per option
noul—noul — the probability of yes
scorescore — the expected level, counting the first as 0; legend maps each level back to your wordingprobabilities — one number per level, keyed "0", "1", …

confidence is one minus the normalized entropy of the probabilities: 1 when all the weight sits on one option, 0 when it is spread evenly. It describes spread, not correctness, so set your threshold on the probabilities. metadata reports the temperatures applied and how many evaluations the call ran; options.temperature in a request overrides the default.

usage.input_tokens counts what you sent. Every answer is a short, fixed-shape value rather than generated text, so there is nothing to meter on the way out. How image inputs are counted is still being confirmed.

Several questions in one call

questions is a map: you choose the keys — route, refund, urgency, or anything else that reads well in your own code — and describe one question under each. The example above asks three questions about the same state and gets all three answers back together in one response, each under the same key it was asked with.

Note

Because every answer is a short, fixed-shape value rather than a generated paragraph, asking three questions in one call costs less than three separate free-text generations would. Exactly how latency scales with each extra question is still being measured — see Limits and pricing.

Probabilities and thresholds

Every answer carries a probability: the model's estimate, given what's in state, that this option is the right one — not a measure of how clearly the question was asked, and not a guess dressed up as a number.

A calibrated model's probabilities mean what they say. Pull out every answer where it said "90% confident" — a well-calibrated model is right in about 90 of every 100 of them. A model that runs overconfident is right less often than its own numbers claim; one that runs underconfident is right more often than it admits.

Your code sets the threshold, not the model. Set the bar higher for decisions that are expensive to get wrong, and lower where a wrong guess costs little; anything below it is a signal to route the case to a person instead of forcing a choice. In the example above, a 60% bar lets your code act on route (88%) and refund (94%) directly, and sends urgency, which came back high at only 58%, to a person.

We report calibration two ways. ECE, expected calibration error: group every answer by the probability the model gave it, in bands (a 0.90–0.95 band, an 0.80–0.85 band, and so on); in each band, compare the model's average stated probability to how often it was actually right; ECE is the size of that gap, averaged across bands and weighted by how many answers fall in each one. Distance from the stated odds: on questions that state the true odds, the total-variation distance between the probabilities the model returns and those odds. Zero is perfect for both; lower is better — see them measured.

How it compares to a general-purpose model

A general-purpose model can do almost anything with language. Standard One does one thing: given a state and a small set of options, it picks. That trade shows up in four places.

A general-purpose modelStandard One
What comes backA paragraph you still have to read or parseOne of the options you gave it
How sure it isUsually described in words, if at allA probability for every option
FormatYou parse the reply and retry when it breaksAlways the shape you asked for
Speed and costGrows with how much it writesOne short answer per question, however many you ask
Note

A typed answer can't come back malformed — that's a guarantee about shape, not about truth. A miscalibrated or simply wrong answer can still arrive with a confident-looking probability attached. Treat the number as a calibrated estimate worth weighing against a threshold, not a promise that it's correct — see Probabilities and thresholds.

Where teams use it

Either size, with the questions you write. A few common shapes:

WorkflowWhat it decidesSample answer
Support triageWhich team should take this tickettechnical · 0.95
Content moderationWhether a photo can be publishedpublish · 0.98
Lead scoringHow warm an inbound request ishot · 0.63
Expense approvalsApprove, reject, or ask for detailsask for details · 0.86
Search relevanceWhether a passage answers the questionno · 0.88
Agent routingWhich tool an agent should call nextcheck_calendar · 0.86
Note

The sample answers show the shape of a reply; they aren't measured results. Two of these shapes are weak without fine-tuning: ten-way support triage and passage relevance for retrieval score 36% and 59% on One 8B. The support triage example on the Standard One page uses four options. For ten-way triage or retrieval relevance, fine-tune the open weights on your own labelled cases.

Call the API

Every request is a POST to /v1/systemone with your state and questions; the response carries one answer per question, in the shape described above. You'll need a Standard Thinking API key — see Create an API key if you don't have one yet.

curl https://api.standardthinking.ai/v1/systemone \
  -H "Authorization: Bearer $STANDARD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "standard-one-3b",
    "state": "I was charged twice for order #48213. Please refund the extra charge today.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Who should pick this up?",
        "criteria": {
          "billing": "Payments, charges, and refunds.",
          "technical": "A bug or a broken feature.",
          "shipping": "A delivery or fulfillment problem.",
          "sales": "A new purchase or an upgrade."
        }
      },
      "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for a refund?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is it?",
        "criteria": [
          "Low: can wait a few days.",
          "Medium: should be handled today.",
          "High: needs attention right now."
        ]
      }
    }
  }'

Set model to standard-one-3b or standard-one-8b; nothing else in the request changes.

Python and JavaScript SDKs, CSV batch processing, and a LangChain integration are planned.

Deployment and data

Both sizes are called the same way as any other Standard Thinking model, over HTTPS, from the regions we publish for the API. Both can also be run from their open weights, linked in Choose a size. A dedicated deployment — capacity reserved for one team rather than shared — is planned for teams that need one.

Note

"Dedicated" can mean several different things — separate capacity, a private network path, or data that never leaves a chosen region. We'll say exactly which of these a dedicated deployment provides once it's confirmed, rather than imply all of them.

Evaluation

We measured both sizes in the release configuration: merged BF16 weights behind SGLang 0.5.20 and jev-adapter, native prompt wording, no system prompt, one option order, and the temperature fitted for each answer type. The wording and the temperatures were chosen on data outside JevBench. Jev 1.13 was measured on the same questions through its hosted endpoint, with its raw probabilities. These are our measurements, not official sealed-set JevBench scores.

Note

These results compare models on one narrow kind of task — picking from a short list of options, under a strict response format — not on writing or conversation generally. A general-purpose model given room to answer in its own words, or wired up with its own constrained decoding, can close some of this gap on an equivalent narrow task. Read the comparison on that basis, not as a general capability ranking.

Accuracy on the served endpoint (higher is better)

Suite (questions)One 3BOne 8BJev 1.13
JevBench public, easy (48)100.00%100.00%100.00%
JevBench public, standard (72)86.11%93.06%98.61%
JevBench public, hard (111)47.75%54.95%72.07%
Judge proxy: routing and answer adequacy (600)87.50%89.33%90.50%
Realistic transfer set (600)90.17%90.83%86.67%
Stated-distribution probability (1,036)78.38%81.18%72.97%
Hard proxy (600)44.17%53.00%54.83%
No JevBench question was used to choose the wording or the temperatures. One 3B is version 2.1 and One 8B version 2.

Latency, p50 (ms, lower is better)

One 3B

22.6 ms

One 8B

25.8 ms

One request at a time on one NVIDIA H200, at about 280 input tokens per decision; p95 is 33.2 ms and 41.9 ms. On the same profile One 8B handles about 40 decisions a second with eight requests in flight. Other request lengths, concurrency and hardware give different figures.

Probability quality (lower is better)

MeasureOne 3BOne 8BJev 1.13
Distance from the stated odds (total variation)0.1240.1080.192
ECE on JevBench public, hard (111)0.2120.1840.099
One 3B and One 8B use their fitted temperatures; Jev 1.13's probabilities are raw. Both measures are defined in Probabilities and thresholds.

Public classification suites (400 cases each, higher is better)

SuiteOne 3BOne 8B
Topics, AG News84.5%84.2%
Typed decisions67.3%71.1%
Intent, MASSIVE, mean of nine languages83.4%86.1%
Email spam92.0%91.8%
Phishing88.8%85.2%
Served endpoint, seed 13. Results vary by language and task.
  • Versions: One 3B v2.1, released September 27, 2026; One 8B v2, released September 26, 2026. One 8B was measured from September 24 to 26, 2026.
  • Every accuracy figure is for text decisions. Accuracy on image inputs hasn't been benchmarked yet.
  • A sealed-set JevBench v1.4.1 run has been requested; its result isn't available yet.
  • Full tables, per-language results, offline runs and the server code are on the model cards: StandardOne-3B and StandardOne-8B.

Limits and pricing

LimitValue
Questions per callUp to 256
Options per choice2 to 26
Levels per score2 to 10
Context8,192 tokens
Image formatData URLs (data:image/…;base64,…) or HTTP(S) URLs, in images
Image sizeTo be published
LanguagesEnglish, Japanese, Chinese, Spanish, French, German, Portuguese and Russian, with a smaller share of Korean; accuracy varies by language
Request rateSee general rate limits
Price, One 3B$0.019 per 1M text input tokens, $0.036 per 1M image input tokens
Price, One 8B$0.049 per 1M text input tokens, $0.084 per 1M image input tokens
Output tokensNot billed on either size

Frequently asked questions

Does Standard One read images?

Yes, both sizes do. Send pictures in images beside a short state — see What you send. The published accuracy figures are for text; accuracy on images hasn't been benchmarked yet.

Can I run Standard One myself?

Yes. Both sizes are open weights, published on Hugging Face as StandardThinking/StandardOne-3B and StandardThinking/StandardOne-8B, and both are also available through the API — see Choose a size. The server code in the StandardOne-8B repository answers on the same path, /v1/systemone, with the same request and response.

What does noul mean?

It's our API's name for a yes-or-no question. It returns noul, the probability that the answer is yes. If the question needs it, say what yes and no mean in criteria under "true" and "false".

Can Standard One write free text?

No. Every answer is one of the types you asked for — a choice, a yes-or-no, or a score — never a paragraph. For open-ended writing or conversation, use a general-purpose model instead.

What happens when it isn't sure?

Your own code decides. Choose the probability above which it acts on its own; below that, route the case to a person instead of forcing a choice — see Probabilities and thresholds.

Is there a live demo?

Yes. The API is live: sign in and send it a request — see Call the API. To see it in use first, watch One 3B decide a support queue, or read three worked examples on the Standard One page. The messages and answers in both are illustrations; measured results are in Evaluation.

Are there SDKs?

Python and JavaScript SDKs, CSV batch processing, and a LangChain integration are planned; package names and availability are being confirmed.