The Responses API is OpenAI's next-generation agentic interface. Compared with Chat Completions, it offers more powerful capabilities:
- Function Calling: the model can call your custom functions
- Built-in tools:
web_search_preview(web search) and more, out of the box - Server-side multi-turn context: use
previous_response_idto maintain conversation history automatically, without the client sending the full message list - Reasoning effort control: precisely tune the depth of thinking via
reasoning.effort
Note: Models marked Responses Only (such as
gpt-5-pro-official,gpt-5.3-codex-official) support only this API, not Chat Completions. For the full model list, see the Model Overview.
Authorizations
Bearer Token authentication
Authorization: Bearer YOUR_API_KEYGet an API Key: visit the API Key management page
Body
Model name
Example: "gpt-5-pro-official", "gpt-5.3-codex-official", "gpt-5.2-official"
User input, in one of two formats:
- String: simple text input
- Message array: multi-turn conversation format
System instructions that guide the model's behavior (equivalent to the system message in Chat Completions)
falseWhether to enable streaming output
Maximum number of tokens to generate
1Sampling temperature, range 0 ~ 2
1Nucleus sampling probability threshold, range 0 ~ 1
ID of the previous response, used to automatically chain multi-turn context server-side without the client sending the full message history
Reasoning configuration
List of available tools
autoTool selection strategy: auto, none, required
Response
Unique identifier for the response (can be used as previous_response_id)
Always response
Response status: completed, failed, in_progress
List of output items, which may include several types:
message: a text reply, containingcontent[].textfunction_call: a function call request, containingnameandargumentsreasoning: the reasoning process (appears whenreasoning.effortis notnone)web_search_call: a record of a web search call
Token usage statistics
usage.input_tokens: number of input tokensusage.output_tokens: number of output tokensusage.output_tokens_details.reasoning_tokens: number of reasoning tokensusage.total_tokens: total number of tokens
Background tasks
Requests submitted with background: true return immediately with status queued / in_progress; poll and cancel afterwards (same API Key as submission; 7-day query window):
GET /v1/responses/{response_id}POST /v1/responses/{response_id}/cancel