Standard API
Open-model inference, billed per token
Call supported models through the OpenAI-compatible Standard API. Input and output tokens are billed separately. Cache-read rates apply under each provider’s own matching rules; how prompt caching is billed on open-model APIs compares them.
Choose a model for your workload.
Reasoning, coding, and multimodal models, all through one API.
Connect your application in three steps.
Set your API key, the Standard Thinking base URL, and a model ID enabled for your workspace. Check supported features in the reference.
1 · Set your API key
Open the console and create a key for your project.
2 · Set the base URL
Set the Standard Thinking base URL and load your API key from a server-side environment variable.
3 · Pick an enabled model ID
Use GET /v1/models to find the model IDs available to your project.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.STANDARD_API_KEY,
baseURL: "https://api.standardthinking.ai/v1",
});
const response = await client.chat.completions.create({
model: "deepseek-v4-1-flash",
messages: [
{ role: "user", content: "Explain this architecture." },
],
stream: true,
});
for await (const chunk of response) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
Request metadata example.
Illustrative fields for usage and troubleshooting. These sample values are not measured latency or a performance claim.
| Time | Model | Tokens in / out | Latency | Content retention | Cost |
|---|---|---|---|---|---|
| 14:02:11 | deepseek-v4-1-flash |
1,204 / 388 | 412 ms | Not retained | $0.0007 |
| 14:02:09 | glm-5-3 |
2,410 / 1,120 | 1.8 s | Not retained | $0.0071 |
| 14:02:09 | deepseek-v4-1-flash |
640 / 212 | 296 ms | Not retained | $0.0004 |
| 14:02:08 | deepseek-v4-1-flash |
3,980 / 1,502 | 2.4 s | Not retained | $0.0024 |
| 14:02:08 | glm-5-3 |
512 / 96 | 640 ms | Not retained | $0.0008 |
Sample request records. Values are illustrative and do not represent a performance benchmark. Retention depends on the selected service and configuration.
API data handling.
Request content, operational metadata, and stored features have different data rules. For how other inference providers state retention and training, read what inference providers keep after a request.
Request content
Retention and temporary caching depend on the service and configuration. ZDR applies where expressly specified.
Operational metadata
Content-free metadata is kept for billing, reliability, and service operations.
Stored features
Features that store or reuse content have separate retention rules and controls.
Model training
Inputs and outputs are not used to train public or shared models without separate, explicit opt-in.
Choose the right access for your work.
Standard API and Product Plan both support products. Coding Plan covers a developer’s own work.
Do I need a Product Plan for production?
Standard API also supports products. Consider a Product Plan for monthly base capacity, free extra capacity when available, and optional auto scaling. Product Plan also includes a data platform where your product’s data is stored under your control.
Can I call every listed model?
Access depends on the model and workspace. Models offered by agreement require pricing and access confirmation before use.
Is this the same as Coding Plan?
Standard API bills requests by token usage, including requests from coding applications. Coding Plan is a monthly plan for a developer’s own work.