Fine-tuned classifier or LLM: accuracy, speed and cost compared
With thousands of examples and a fixed label set, a fine-tuned encoder still wins on accuracy. Without them, or when labels change or need a probability, a decision model starts the same day at a similar cost per request.
A fine-tuned text classifier is still the most accurate way to sort text into a fixed set of labels when you have thousands of examples of each, and one of the cheapest to run. A model that returns a probability for each of your labels, with no training step on your side, fits the other cases: no examples yet, labels that change, several languages, or a need for a number to set a review threshold. Standard Thinking serves one such decision model, Standard One. The independent figures below come from two 2026 studies, and the Standard One figures from its model cards.
Three ways to label the same text
| Fine-tuned encoder | LLM prompt | Decision model | |
|---|---|---|---|
| Labelled data first | Thousands | None | None |
| Adding a label | Retrain, redeploy | Edit the prompt | Edit the options |
| What comes back | A label and scores | Text to parse | A probability per option |
| AG News | 94.6 macro-F1 | 91.4 macro-F1 | 84.5% / 84.2% |
| Median latency | 108–197 ms | 332–1,435 ms | 22.6 / 25.8 ms |
| Per million requests | $5.73–$10.44 | $276–$1,272 | $5.32 / $13.72 |
- Encoder and LLM prompt columns come from Valdes Gonzalez (2026). DistilBERT, BERT and RoBERTa were fine-tuned on AG News and served on CPU through Google Cloud Run; GPT-4o and Claude 4.5 were prompted zero- and few-shot through their APIs; all were scored on the 3,800-item test set, with latency measured end to end.
- AG News shows the best encoder (RoBERTa) and the best prompt (Claude 4.5 zero-shot) in that study as macro-F1. The decision-model figures are Standard One 3B and 8B accuracy on 400 cases sampled from the same test split; AG News is not among the public datasets in Standard One's training mixture. The two metrics are close on this balanced dataset but not identical, and the samples differ.
- Median latency for the decision model is p50 at the model server on one H200, for 242 decisions of about 280 input tokens run one at a time. Network time to the caller comes on top, so it does not compare like for like with end-to-end figures.
- Per million requests is hosting for the encoders and token charges for the prompts. For the decision model it is 280 input tokens per decision at Standard One's current rates of $0.019 (3B) and $0.049 (8B) per million text input tokens, half the list price, with output free.
When to train the classifier
- A fixed label set with plenty of examples. On AG News the tuned RoBERTa leads the best prompt by about three points and Standard One by about ten, and the study finds tuned encoders "Pareto-optimal across most realistic operating regimes".
- Intents with abundant in-domain data. Rodrigues and Vas (2026) measured a fine-tuned RoBERTa at 95.9% on ATIS against 84.1% for a zero-shot LLM, at three orders of magnitude lower cost.
- A task the decision model is weak at without tuning. Standard One's 8B card reports ten-way support triage at 36% and RAG passage relevance at 59%, and advises fine-tuning for those.
When a decision model is the better start
- No labels yet. It answers from the option descriptions you write, and the cases reviewers correct become a training set if you later move a label to a classifier.
- Labels that change. Rodrigues and Vas report that a classifier trained on one app's intents scored 0% on a new app's intents, while a schema-prompted LLM served both at about 94% with no retraining. A decision model reads its options from each request, so a new label is an edit, not a training run.
- Several languages. One model serves the eight languages of its training mixture and a smaller Korean share. On MASSIVE intent, with 20 options in each of nine languages, Standard One averages 83.4% (3B) and 86.1% (8B); MASSIVE's training split is part of its training mixture, so these are not zero-shot figures.
- Out-of-scope requests. In the same study the LLM recalled 85.6% of out-of-scope requests, against 58.1% for the tuned RoBERTa.
- A number to act on. Every option comes back with a probability, which is what a review threshold needs: see how to set an LLM confidence threshold.
Start with one, move labels to the other
The two can share a pipeline. Ship with the decision model and a review threshold, keep the labels reviewers assign, and train a classifier for a label set once it has settled and gathered a few thousand examples. Before moving a label over, compare both on the same held-out sample, and keep the decision model for new labels, other languages and out-of-scope traffic.
Standard One 3B and 8B are open weights under Apache 2.0, served through POST /v1/systemone: see the model page and the documentation.
Sources and verification
Sources used for this resource. Verification dates describe our source checks, not provider effective dates.
Valdes Gonzalez, Cost-Aware Model Selection for Text Classification (arXiv 2602.06370)
Version 1, 6 February 2026, read in full text. AG News: "All three encoders cluster around ≈94–95 macro-F1 (best: RoBERTa 94.63 ± 0.14), whereas LLM prompting lags substantially even with few-shot (best LLM: Claude 4.5 ZS 91.35 ± 0.11; GPT-4o FS 89.65 ± 0.13)"; "encoder p50 latency is 108–197 ms (p95 161–286 ms), while LLM p50 latency is 332–1435 ms"; "Cost at scale strongly favors encoders: $5.73–$10.44 per 1M requests versus $276.00–$1271.58 for LLM prompting." Encoders were "deployed as stateless inference services on Google Cloud Run" on CPU; AG News has 120,000 training and 3,800 test items. The conclusion quoted: "encoder-based models remain Pareto-optimal across most realistic operating regimes".
Rodrigues and Vas, When Do LLMs Replace Fine-Tuned NLU? (arXiv 2608.20371)
Submitted 19 June 2026; the LLM is Claude Haiku zero-shot. From the abstract: on ATIS fine-tuned RoBERTa "beats Claude zero-shot by 11.8 points (95.9 vs. 84.1, p<0.001)" and is "three orders of magnitude cheaper and faster"; on CLINC150 the two tie (89.1 vs. 88.5); "out-of-scope detection (OOS recall 85.6 vs. 58.1 for RoBERTa)"; and "a classifier trained on one app's intents scores 0% on a new app's intents while the schema-prompted LLM serves both at ~94% with zero retraining".
Standard One 3B model card (v2.1)
Public classification and decision suites, v2.1: "Served endpoint, 400 cases per suite, seed 13: AG News 84.5 %, typed decisions 67.3 %, nine-language MASSIVE intent mean 83.4 %, email spam 92.0 %, phishing 88.8 %." Training languages: "English, Japanese, Chinese, Spanish, French, German, Portuguese, Russian and a smaller Korean share."
Standard One 8B model card (v2)
Public suites (400 cases/suite, seed 13, served endpoints): "AG News 84.2 %, typed decisions 71.1 %, MASSIVE intent mean 86.1 %, email spam 91.8 %, phishing 85.2 %". Latency: raw serial p50 22.6 ms (3B) and 25.8 ms (8B) on one H200 with SGLang 0.5.20 for a 242-decision profile averaging about 280 input tokens, "not a latency guarantee for other request lengths, concurrency or hardware". Weak workflows: "Ten-way support triage (36 %) and RAG passage relevance (59 %) are weak zero-shot; fine-tune for those."
Standard One public classification suites
Suite definitions only: AG News is ag_news, test split, 4 labels, 400 cases, seed 13; MASSIVE intent has 20 options per language in nine languages (en, ja, zh, es, fr, de, pt, ru, ko). The tables in this file are historical v1.1 figures and are not quoted.
Standard One benchmark report, training data provenance
The public datasets converted into training items (train splits) include MASSIVE (AmazonScience/massive), CLINC150, Banking77 and DBpedia-14, and do not include AG News; the report states that its provenance sections already describe the v2 weights. This is the basis for calling the MASSIVE figures in-distribution and the AG News figures free of task training.
Standard Thinking pricing, Standard One
The rates used for the per-decision cost: text input at $0.019 per 1M tokens for One 3B and $0.049 for One 8B, each shown as 50% off list prices of $0.038 and $0.098, and the sentence "Choose a size and pay per input token. Each call returns one short answer, so output is free."
Continue exploring
2.6%worst case when 3 in 300 are wrong
- Observed
- 1.0%
- 8B ECE
- 0.184
- 3B ECE
- 0.212
95% Clopper–Pearson bound; ECE from the Standard One model cards
How to set a confidence threshold for automated LLM decisions
A threshold decides which answers your code acts on alone. Set it on the chosen answer's probability, from a labelled sample of your own traffic, and measure again when the traffic changes.
19.6Mtokens a month
- Input
- 18.8M
- Output
- 850K
- Model calls
- 5K
Simple tool loop · 1,000 runs a month
Plan tokens for chats and agents
Account for repeated conversation history and tool calls before choosing a rate card.
1,000 calls a day · 30 days · 50% cached input
What will GPT-OSS 120B cost?
Compare four rate cards, including ours, using your request volume, token usage, and cache assumptions.