Responses API

The OpenAI Responses API format, with function calling, built-in tools, and server-side multi-turn context management

POST/v1/responses

The Responses API is OpenAI's next-generation agentic interface. Compared with Chat Completions, it offers more powerful capabilities:

  • Function Calling: the model can call your custom functions
  • Built-in tools: web_search_preview (web search) and more, out of the box
  • Server-side multi-turn context: use previous_response_id to maintain conversation history automatically, without the client sending the full message list
  • Reasoning effort control: precisely tune the depth of thinking via reasoning.effort

Note: Models marked Responses Only (such as gpt-5-pro-official, gpt-5.3-codex-official) support only this API, not Chat Completions. For the full model list, see the Model Overview.

Authorizations

AuthorizationstringRequired

Bearer Token authentication

Authorization: Bearer YOUR_API_KEY

Get an API Key: visit the API Key management page

Body

modelstringRequired

Model name

Example: "gpt-5-pro-official", "gpt-5.3-codex-official", "gpt-5.2-official"

inputstring | object[]Required

User input, in one of two formats:

  • String: simple text input
  • Message array: multi-turn conversation format
ShowHide Message array format
rolestringRequired

Message role: user, assistant, developer

contentstring | object[]Required

Message content, plain text or an array of content blocks including images

instructionsstring

System instructions that guide the model's behavior (equivalent to the system message in Chat Completions)

streambooleanDefault false

Whether to enable streaming output

max_output_tokensinteger

Maximum number of tokens to generate

temperaturenumberDefault 1

Sampling temperature, range 0 ~ 2

top_pnumberDefault 1

Nucleus sampling probability threshold, range 0 ~ 1

previous_response_idstring

ID of the previous response, used to automatically chain multi-turn context server-side without the client sending the full message history

reasoningobject

Reasoning configuration

ShowHide reasoning
effortstring

Reasoning effort: high, medium, low, none

Higher values make the model think more deeply and consume more reasoning tokens

toolsobject[]

List of available tools

ShowHide tools[n]
typestringRequired

Tool type: function (custom function), web_search_preview (web search)

namestring

Function name (required when type is function)

descriptionstring

Function description

parametersobject

JSON Schema for the function parameters

tool_choicestringDefault auto

Tool selection strategy: auto, none, required

Response

idstring

Unique identifier for the response (can be used as previous_response_id)

objectstring

Always response

statusstring

Response status: completed, failed, in_progress

outputobject[]

List of output items, which may include several types:

  • message: a text reply, containing content[].text
  • function_call: a function call request, containing name and arguments
  • reasoning: the reasoning process (appears when reasoning.effort is not none)
  • web_search_call: a record of a web search call
usageobject

Token usage statistics

  • usage.input_tokens: number of input tokens
  • usage.output_tokens: number of output tokens
  • usage.output_tokens_details.reasoning_tokens: number of reasoning tokens
  • usage.total_tokens: total number of tokens

Background tasks

Requests submitted with background: true return immediately with status queued / in_progress; poll and cancel afterwards (same API Key as submission; 7-day query window):

GET  /v1/responses/{response_id}POST /v1/responses/{response_id}/cancel