It sounds like the carrier may have left order #57120 at the wrong address, since the proof-of-delivery photo doesn’t match. The customer could ask neighbours and the carrier, and it’s also worth checking the address saved on the account. A replacement or a refund may be appropriate. Given the Friday flight, this seems fairly time-sensitive. If you’d like, I can draft a reply to the customer.
Standard One
Show it a message.
Get a decision back.
A decision model in two sizes that reads messages and pictures, for the calls your product makes all day: which team, which label, yes or no, how urgent. Each size answers with one of your own options and says how sure it is.
Decide what it handles on its own.
Every answer comes with a probability. Set a bar: your code acts above it and hands the rest to a person.
Not a chatbot. An answer your code can use.
A general-purpose model writes a paragraph that someone, or some code, still has to read. Standard One returns one of your options and how sure it is.
{
"answers": {
"route": { "type": "choice", "probabilities": {
"shipping": 0.94, "technical": 0.03, "billing": 0.02, "sales": 0.01 },
"confidence": 0.79, "choice": "shipping" },
"reship": { "type": "noul", "noul": 0.86 },
"urgency": { "type": "score", "probabilities": {
"0": 0.06, "1": 0.23, "2": 0.71 },
"confidence": 0.32, "score": 1.65,
"legend": { "0": "low", "1": "medium", "2": "high" } }
},
"usage": { "input_tokens": 61, "output_tokens": 0 }
}
Pick the size that fits each decision.
Both take the same request, read text and images, and return the same kind of answer.
One 3B
standard-one-3bFaster, for high-volume calls.
- Realistic decisions
- 90.2%
- Median latency
- 22.6 ms
- Price
- $0.019 text · $0.036 image per 1M input tokens, 50% off list
- Inputs
- Text and images
- Availability
- API · Open weights (Apache 2.0)
Best for spam and phishing filters, topic labels and other high-volume checks.
One 8B
standard-one-8bStronger on hard cases, for high-stakes calls.
- Realistic decisions
- 90.8%
- Median latency
- 25.8 ms
- Price
- $0.049 text · $0.084 image per 1M input tokens, 50% off list
- Inputs
- Text and images
- Availability
- API · Open weights (Apache 2.0)
Best for agent routing, answer checks and intent across languages.
Anywhere your product has to make a call.
Support queues, back offices and agents use it the same way: you write the questions and set the bar. Pick One 8B where a wrong call costs the most and One 3B where speed matters most.
Support triage
Every message in the inbox goes to One 3B with one question and the four teams as the options. It answers fast enough to run on each message as it arrives.
Above 0.85 your code routes the message. Below it, an agent picks the team.
IIbrahim Sow9:12
The export button on the Reports page has been greyed out since yesterday’s release.
Which team should handle this?
Technical0.95
Sales0.02
Billing0.02
Shipping0.01
Routed to Technical
GGrace Oyelaran9:14
The app wiped my saved addresses and my order shipped to the old one.
Which team should handle this?
Technical0.48
Shipping0.46
Billing0.04
Sales0.02
Held for an agent
Expense approvals
Each claim goes to One 8B with the receipt and the amount claimed, and three options: approve, reject or ask for details. A wrong approval costs real money, so the larger size takes this one.
Above 0.80 your code acts on the answer. Below it, finance looks at the claim.
RRavi Menon
- Receipt total
- $210.72
- Claimed
- $230.72
Approve, reject or ask for details?
Ask for details0.86
Approve0.08
Reject0.06
Asks Ravi about the $20.00 difference
Agent routing
Before each step the agent asks One 8B which tool to run next, with its tools as the options. Here the request hangs on a meeting time, so the calendar comes first.
Above 0.80 the agent calls the tool it picked. Below it, the agent checks with the user first.
HHana Sato18:40
Find flights to Lisbon next Friday after my 4 pm meeting.
Which tool should run next?
check_calendar0.86
search_flights0.11
ask_user0.03
Calls check_calendar
How Standard One works
The request format, the three kinds of questions, probabilities and thresholds, and how we measured speed, accuracy and calibration.