Standard One

Show it a message.
Get a decision back.

A decision model in two sizes that reads messages and pictures, for the calls your product makes all day: which team, which label, yes or no, how urgent. Each size answers with one of your own options and says how sure it is.

standard-one-3b Sample answer

Decide what it handles on its own.

Every answer comes with a probability. Set a bar: your code acts above it and hands the rest to a person.

Hollow dots are the answers it acted on and got wrong.
Handled on its own
81%
Sent to a person
19%
Right when it acts
97%

An illustration with 120 simulated decisions, not a measurement. How thresholds work.

Not a chatbot. An answer your code can use.

A general-purpose model writes a paragraph that someone, or some code, still has to read. Standard One returns one of your options and how sure it is.

Noor Haddad10:31Sample

Delivered to the wrong porch

Tracking says my order #57120 was delivered at 10:14, but the photo shows a porch that isn’t mine. I need it before I fly on Friday.

General-purpose model3.4 s

It sounds like the carrier may have left order #57120 at the wrong address, since the proof-of-delivery photo doesn’t match. The customer could ask neighbours and the carrier, and it’s also worth checking the address saved on the account. A replacement or a refund may be appropriate. Given the Friday flight, this seems fairly time-sensitive. If you’d like, I can draft a reply to the customer.

…and your code still has to parse it.

One 3B
22.6 ms
RouteShipping 0.94
ReshipYes 0.86
UrgencyHigh 0.71

Pick the size that fits each decision.

Both take the same request, read text and images, and return the same kind of answer.

One 3B

standard-one-3b

Faster, for high-volume calls.

Realistic decisions
90.2%
Median latency
22.6 ms
Price
$0.019 text · $0.036 image per 1M input tokens, 50% off list
Inputs
Text and images
Availability
API · Open weights (Apache 2.0)

Best for spam and phishing filters, topic labels and other high-volume checks.

Weights on Hugging Face

One 8B

standard-one-8b

Stronger on hard cases, for high-stakes calls.

Realistic decisions
90.8%
Median latency
25.8 ms
Price
$0.049 text · $0.084 image per 1M input tokens, 50% off list
Inputs
Text and images
Availability
API · Open weights (Apache 2.0)

Best for agent routing, answer checks and intent across languages.

Weights on Hugging Face

Accuracy on 600 realistic decisions. Latency is the median on one NVIDIA H200, at about 280 input tokens per decision.

Anywhere your product has to make a call.

Support queues, back offices and agents use it the same way: you write the questions and set the bar. Pick One 8B where a wrong call costs the most and One 3B where speed matters most.

Support triage

One 3B4 optionsBar 0.85

Every message in the inbox goes to One 3B with one question and the four teams as the options. It answers fast enough to run on each message as it arrives.

Above 0.85 your code routes the message. Below it, an agent picks the team.

Expense approvals

One 8B3 optionsBar 0.80

Each claim goes to One 8B with the receipt and the amount claimed, and three options: approve, reject or ask for details. A wrong approval costs real money, so the larger size takes this one.

Above 0.80 your code acts on the answer. Below it, finance looks at the claim.

Agent routing

One 8B3 toolsBar 0.80

Before each step the agent asks One 8B which tool to run next, with its tools as the options. Here the request hangs on a meeting time, so the calendar comes first.

Above 0.80 the agent calls the tool it picked. Below it, the agent checks with the user first.

Sample messages and answers, written for illustration. Measured results are in the evaluation.

Technical write-up · 12 min read

How Standard One works

The request format, the three kinds of questions, probabilities and thresholds, and how we measured speed, accuracy and calibration.

Question typesRequest formatThresholdsEvaluation
Read the write-up

Distance from the stated odds · lower is better

One 3B

0.124

One 8B

0.108

On questions that state the true odds, this is the average distance between the probabilities it returns and those odds. Zero would be an exact match.