Skip to content

Connect to Standard Thinking.

Explore Standard One, set up the Standard API, and find access and integration guidance for your plan.

API reference: /v1 · Documentation updated

Send your first API request.

You need an account, an API key, and access to the model used in the example. Follow the installation instructions for your language before running the code.

1. Create an API key

Open the console and create a key for your project. Copy the secret when it is shown and store it securely.

Keep your key private

Keep API keys out of browser code and source control. Use a separate key for each service or environment.

Set your API key
export STANDARD_API_KEY="YOUR_STANDARD_API_KEY"

2. Configure your client

Set the Standard Thinking base URL and load your API key from a server-side environment variable.

Base URL
https://api.standardthinking.ai/v1

3. Check model access

Use GET /v1/models to find the model IDs available to your project. Compare commercial access requirements in the model catalog.

GET200 · JSON
/v1/models
curl
curl https://api.standardthinking.ai/v1/models \
  -H "Authorization: Bearer $STANDARD_API_KEY"

4. Send and read a response

Choose an example below and use a model enabled for your workspace. The JavaScript and Python examples display the streamed text; curl displays the raw streaming events. These examples use deepseek-v4-1-flash.

Install dependencies

JavaScript: run npm install openai, save the example as request.mjs, then run node request.mjs. Python: run python3 -m pip install openai, save as request.py, then run python3 request.py. Set the API key in the same terminal first.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.STANDARD_API_KEY,
  baseURL: "https://api.standardthinking.ai/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek-v4-1-flash",
  messages: [{ role: "user", content: "Explain this architecture." }],
  stream: true,
});

for await (const chunk of response) {
  process.stdout.write(chunk.choices[0]?.delta?.content || "");
}

API reference

Check supported requests, response fields, and error handling. Model access, feature support, and limits can vary by workspace and model.

Authentication

Authenticate every request with a project API key in the Authorization header.

HeaderRequired
Authorization: Bearer $STANDARD_API_KEY
Key scope

Keys inherit project limits and can be individually revoked. Use a separate key for each production service.

Chat completions

POSTJSON
/v1/chat/completions

Send a conversation to a supported model. Check its supported inputs and parameters before making a request. Set stream to true for incremental server-sent events.

Request parameters

Parameter Type Description
model string · required The model ID to use for generation.
messages array · required Conversation messages in system, user and assistant roles.
stream boolean Returns incremental SSE events when true.
max_tokens integer Maximum number of output tokens to generate.
temperature number Sampling temperature between 0 and 2.
tools array Functions the model may call during generation.

Response

Field Type Description
id string Unique identifier for the completion.
choices[] array Generated messages or tool calls, with a finish reason for each choice.
usage.prompt_tokens integer Input tokens metered for the request.
usage.completion_tokens integer Output tokens metered for the response.

Cached-input billing and the usage fields that report it differ by provider; how prompt caching is billed on open-model APIs compares them.

Streaming

Streaming responses use server-sent events. Each event carries a partial delta, followed by a final [DONE] marker.

data: {"choices":[{"delta":{"content":"Standard Thinking"}}]}

data: {"choices":[{"delta":{"content":" is ready."}}]}

data: [DONE]

Models

GETJSON
/v1/models

List the models available to your project. For commercial access and pricing, see the model catalog and API rates.

Rate limits

Check the request and token limits that apply to your project and model. Use the documented response headers when handling a limit. Your workspace limits are managed in the console.

Header Type Description
x-ratelimit-limit-requests header Maximum requests in the current window.
x-ratelimit-remaining-requests header Requests available before reset.
x-ratelimit-reset-requests header Time until the request window resets.
x-ratelimit-limit-tokens header Maximum tokens in the current window.
x-ratelimit-remaining-tokens header Tokens available before reset.

Error codes

Use the status code, error details, and request ID to investigate a failed request.

Code Key Description
400 bad_request The request body or a parameter value is invalid.
401 unauthorized The API key is missing, invalid or revoked.
429 rate_limit The project exceeded its request or token limit. Back off and retry after the window in x-ratelimit-reset-requests.
500 server_error An unexpected error occurred. Retry after a delay. If it persists, contact support with the request ID.

Other products

Coding Plan limits are below. For a Product Plan, plan an integration or ask what applies to your product.

Coding Plan limits and peak-time priority

Every plan includes Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash. Prices are per developer, per month.

PlanUsage allowanceTime limitsPeak-time priority
$50Smaller allowance5-hour and weekly limitsLowest
$100Larger allowanceWeekly limit only; no 5-hour limitAbove $50, below $200
$200No total usage capNo 5-hour or weekly limitsHighest

Peak-time restrictions apply to every plan, including $200. Requests may slow, queue, or time out. Priority is $200 → $100 → $50. An uncapped usage allowance does not guarantee processing speed or concurrency.

Tools that speak OpenAI chat completions connect with the base URL above. Codex CLI speaks only the Responses API: run a proxy that accepts POST /v1/responses, forwards chat completions with streaming and tool calls intact, and returns usage, then point a model_providers entry in ~/.codex/config.toml at it with wire_api = "responses". Claude Code speaks the Anthropic Messages API; ask about your workflow for setup.

Exact allowances and request and concurrency limits are confirmed before activation. Ask about tool support and setup or compare Coding Plan pricing. The vocabulary of allowances, windows, and priority is defined in capacity and plan terms.

If the first request fails

  • Key not accepted: check that the key is valid and loaded in the environment where the request runs.
  • Model unavailable: check the model ID and your workspace’s access to that model.
  • Request limit reached: follow the documented retry timing for the applicable request or token limit.

For help, contact technical support with the request ID and error details. Do not include your API key.

Review streaming and error handling, compare models, and check API rates before connecting your application.