Connect to Standard Thinking.
Explore Standard One, set up the Standard API, and find access and integration guidance for your plan.
Standard One
Compare the two sizes, what they read and return, and how to call them.
Read the Standard One write-upAPI quickstart
Create a key and send a model request.
Start with APIProduct Plan integration planning
Review the workload, base allocation, and service details for an integration.
Plan an integrationCoding Plan access and setup
Compare plan limits and the included Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash models.
Compare Coding Plan limitsStandard One
Technical write-upSend your first API request.
You need an account, an API key, and access to the model used in the example. Follow the installation instructions for your language before running the code.
1. Create an API key
Open the console and create a key for your project. Copy the secret when it is shown and store it securely.
Keep API keys out of browser code and source control. Use a separate key for each service or environment.
export STANDARD_API_KEY="YOUR_STANDARD_API_KEY"
2. Configure your client
Set the Standard Thinking base URL and load your API key from a server-side environment variable.
https://api.standardthinking.ai/v1
3. Check model access
Use GET /v1/models to find the model IDs available to your project. Compare commercial access requirements in the model catalog.
/v1/models
curl https://api.standardthinking.ai/v1/models \
-H "Authorization: Bearer $STANDARD_API_KEY"
4. Send and read a response
Choose an example below and use a model enabled for your workspace. The JavaScript and Python examples display the streamed text; curl displays the raw streaming events. These examples use deepseek-v4-1-flash.
JavaScript: run npm install openai, save the example as request.mjs, then run node request.mjs. Python: run python3 -m pip install openai, save as request.py, then run python3 request.py. Set the API key in the same terminal first.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.STANDARD_API_KEY,
baseURL: "https://api.standardthinking.ai/v1",
});
const response = await client.chat.completions.create({
model: "deepseek-v4-1-flash",
messages: [{ role: "user", content: "Explain this architecture." }],
stream: true,
});
for await (const chunk of response) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
API reference
Check supported requests, response fields, and error handling. Model access, feature support, and limits can vary by workspace and model.
Authentication
Authenticate every request with a project API key in the Authorization header.
Authorization: Bearer $STANDARD_API_KEY
Keys inherit project limits and can be individually revoked. Use a separate key for each production service.
Chat completions
/v1/chat/completions
Send a conversation to a supported model. Check its supported inputs and parameters before making a request. Set stream to true for incremental server-sent events.
Request parameters
| Parameter | Type | Description |
|---|---|---|
| model | string · required | The model ID to use for generation. |
| messages | array · required | Conversation messages in system, user and assistant roles. |
| stream | boolean | Returns incremental SSE events when true. |
| max_tokens | integer | Maximum number of output tokens to generate. |
| temperature | number | Sampling temperature between 0 and 2. |
| tools | array | Functions the model may call during generation. |
Response
| Field | Type | Description |
|---|---|---|
| id | string | Unique identifier for the completion. |
| choices[] | array | Generated messages or tool calls, with a finish reason for each choice. |
| usage.prompt_tokens | integer | Input tokens metered for the request. |
| usage.completion_tokens | integer | Output tokens metered for the response. |
Cached-input billing and the usage fields that report it differ by provider; how prompt caching is billed on open-model APIs compares them.
Streaming
Streaming responses use server-sent events. Each event carries a partial delta, followed by a final [DONE] marker.
data: {"choices":[{"delta":{"content":"Standard Thinking"}}]}
data: {"choices":[{"delta":{"content":" is ready."}}]}
data: [DONE]
Models
/v1/models
List the models available to your project. For commercial access and pricing, see the model catalog and API rates.
Rate limits
Check the request and token limits that apply to your project and model. Use the documented response headers when handling a limit. Your workspace limits are managed in the console.
| Header | Type | Description |
|---|---|---|
| x-ratelimit-limit-requests | header | Maximum requests in the current window. |
| x-ratelimit-remaining-requests | header | Requests available before reset. |
| x-ratelimit-reset-requests | header | Time until the request window resets. |
| x-ratelimit-limit-tokens | header | Maximum tokens in the current window. |
| x-ratelimit-remaining-tokens | header | Tokens available before reset. |
Error codes
Use the status code, error details, and request ID to investigate a failed request.
| Code | Key | Description |
|---|---|---|
| 400 | bad_request | The request body or a parameter value is invalid. |
| 401 | unauthorized | The API key is missing, invalid or revoked. |
| 429 | rate_limit | The project exceeded its request or token limit. Back off and retry after the window in x-ratelimit-reset-requests. |
| 500 | server_error | An unexpected error occurred. Retry after a delay. If it persists, contact support with the request ID. |
Other products
Coding Plan limits are below. For a Product Plan, plan an integration or ask what applies to your product.
Coding Plan limits and peak-time priority
Every plan includes Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Prices are per developer, per month.
| Plan | Usage allowance | Time limits | Peak-time priority |
|---|---|---|---|
| $50 | Smaller allowance | 5-hour and weekly limits | Lowest |
| $100 | Larger allowance | Weekly limit only; no 5-hour limit | Above $50, below $200 |
| $200 | No total usage cap | No 5-hour or weekly limits | Highest |
Peak-time restrictions apply to every plan, including $200. Requests may slow, queue, or time out. Priority is $200 → $100 → $50. An uncapped usage allowance does not guarantee processing speed or concurrency.
Tools that speak OpenAI chat completions connect with the base URL above. Codex CLI speaks only the Responses API: run a proxy that accepts POST /v1/responses, forwards chat completions with streaming and tool calls intact, and returns usage, then point a model_providers entry in ~/.codex/config.toml at it with wire_api = "responses". Claude Code speaks the Anthropic Messages API; ask about your workflow for setup.
Exact allowances and request and concurrency limits are confirmed before activation. Ask about tool support and setup or compare Coding Plan pricing. The vocabulary of allowances, windows, and priority is defined in capacity and plan terms.
If the first request fails
- Key not accepted: check that the key is valid and loaded in the environment where the request runs.
- Model unavailable: check the model ID and your workspace’s access to that model.
- Request limit reached: follow the documented retry timing for the applicable request or token limit.
For help, contact technical support with the request ID and error details. Do not include your API key.
Review streaming and error handling, compare models, and check API rates before connecting your application.