Standard API

Open-model inference, billed per token

Call supported models through the OpenAI-compatible Standard API. Input and output tokens are billed separately. Cache-read rates apply under each provider’s own matching rules; how prompt caching is billed on open-model APIs compares them.

Choose a model for your workload.

Reasoning, coding, and multimodal models, all through one API.

Connect your application in three steps.

Set your API key, the Standard Thinking base URL, and a model ID enabled for your workspace. Check supported features in the reference.

1 · Set your API key

Open the console and create a key for your project.

2 · Set the base URL

Set the Standard Thinking base URL and load your API key from a server-side environment variable.

3 · Pick an enabled model ID

Use GET /v1/models to find the model IDs available to your project.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.STANDARD_API_KEY,
  baseURL: "https://api.standardthinking.ai/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek-v4-1-flash",
  messages: [
    { role: "user", content: "Explain this architecture." },
  ],
  stream: true,
});

for await (const chunk of response) {
  process.stdout.write(chunk.choices[0]?.delta?.content || "");
}

Request metadata example.

Illustrative fields for usage and troubleshooting. These sample values are not measured latency or a performance claim.

Example request log for a sample workspace with ZDR enabled.
Time Model Tokens in / out Latency Content retention Cost
14:02:11 deepseek-v4-1-flash 1,204 / 388 412 ms Not retained $0.0007
14:02:09 glm-5-3 2,410 / 1,120 1.8 s Not retained $0.0071
14:02:09 deepseek-v4-1-flash 640 / 212 296 ms Not retained $0.0004
14:02:08 deepseek-v4-1-flash 3,980 / 1,502 2.4 s Not retained $0.0024
14:02:08 glm-5-3 512 / 96 640 ms Not retained $0.0008

Sample request records. Values are illustrative and do not represent a performance benchmark. Retention depends on the selected service and configuration.

API data handling.

Request content, operational metadata, and stored features have different data rules. For how other inference providers state retention and training, read what inference providers keep after a request.

Request content

Retention and temporary caching depend on the service and configuration. ZDR applies where expressly specified.

Operational metadata

Content-free metadata is kept for billing, reliability, and service operations.

Stored features

Features that store or reuse content have separate retention rules and controls.

Model training

Inputs and outputs are not used to train public or shared models without separate, explicit opt-in.

Choose the right access for your work.

Standard API and Product Plan both support products. Coding Plan covers a developer’s own work.

Do I need a Product Plan for production?

Standard API also supports products. Consider a Product Plan for monthly base capacity, free extra capacity when available, and optional auto scaling. Product Plan also includes a data platform where your product’s data is stored under your control.

Can I call every listed model?

Access depends on the model and workspace. Models offered by agreement require pricing and access confirmation before use.

Is this the same as Coding Plan?

Standard API bills requests by token usage, including requests from coding applications. Coding Plan is a monthly plan for a developer’s own work.