Supported Text Models

An overview of text conversation models, supporting both the Chat Completions and Responses API formats

The platform offers two ways to call the API, switching models via the model parameter:

API format Endpoint Notes
Chat Completions (default) POST /v1/chat/completions Classic conversation endpoint, compatible with all OpenAI SDKs. View docs
Responses (new) POST /v1/responses OpenAI's next-generation Agentic API, supporting function calling, built-in tools, and server-side multi-turn context. View docs

Tip: Models marked Responses Only support only the Responses API, not Chat Completions. All other models are called through Chat Completions by default.


Anthropic / Claude

Model ID Description Vision Extended thinking
claude-haiku-4-5 The fastest, lowest-cost Claude model, ideal for high-concurrency lightweight tasks ✓
claude-sonnet-4-6 The best balance of performance and cost, suitable for most business scenarios ✓
claude-opus-4-6 Claude's flagship model, suitable for complex reasoning and high-quality creation ✓

OpenAI / GPT-5

GPT-5 Official (direct official connection)

Connected directly through the Azure OpenAI official API, offering the highest stability and compliance.

Model ID Description Vision Extended thinking API format
gpt-5-pro-official GPT-5 Pro direct official connection, OpenAI's most powerful flagship reasoning model ✓ ✓ Responses Only
gpt-5.2-official GPT-5.2 direct official connection, multimodal enhanced edition ✓ ✓ All
gpt-5.3-codex-official GPT-5.3 Codex direct official connection, the latest code-optimized edition, supporting function calling and built-in tools ✓ ✓ Responses Only

GPT-5 Standard

Model ID Description Vision Extended thinking
gpt-5 OpenAI's flagship model, powerful general reasoning and multimodal capabilities ✓
gpt-5-codex A GPT-5 optimized for code, excelling at code generation and debugging
gpt-5-codex-mini A lightweight code model with ultra-fast responses at low cost
gpt-5.1 An enhanced GPT-5 with improved instruction following and reasoning ✓
gpt-5.1-codex The code-specialized edition of GPT-5.1
gpt-5.1-codex-max The flagship code edition of GPT-5.1, with the strongest coding capabilities
gpt-5.1-codex-mini A lightweight code edition of GPT-5.1, fast and low-cost
gpt-5.2 The second generation of GPT-5, multimodal enhanced ✓
gpt-5.2-codex The code-specialized edition of GPT-5.2
gpt-5.3-codex The latest code edition of GPT-5, with continuously improved coding capabilities
gpt-5.3-codex-spark The GPT-5.3 Codex Spark variant, focused on high-speed code generation
gpt-5.4 The fourth generation of GPT-5, with stronger general reasoning and multimodal capabilities ✓
gpt-5.5 GPT-5.5, with stronger general reasoning, coding, and multimodal capabilities ✓
gpt-5.6-sol The GPT-5.6 flagship model, suitable for complex reasoning, high-value code, and long-context tasks ✓ ✓
gpt-5.6-terra The GPT-5.6 balanced model, suitable for general production traffic and agent workflows ✓ ✓
gpt-5.6-luna The GPT-5.6 high-throughput model, suitable for classification, extraction, routing, and batch processing ✓ ✓

Google / Gemini

Gemini 3 series

Model ID Description Vision Extended thinking
gemini-3.1-pro-preview-official Gemini 3.1 Pro, the latest-generation flagship reasoning model ✓ ✓
gemini-3.1-flash-lite-preview-official Gemini 3.1 Flash Lite, the most cost-effective lightweight model, with ultra-fast responses ✓
gemini-3-pro-official Gemini 3 Pro flagship, the strongest reasoning and multimodal capabilities ✓ ✓
gemini-3-pro-preview-official The Gemini 3 Pro Preview edition (same as gemini-3-pro-official) ✓ ✓
gemini-3-flash-official Gemini 3 Flash, fast and efficient, a great value choice ✓ ✓
gemini-3-flash-preview-official The Gemini 3 Flash Preview edition (same as gemini-3-flash-official) ✓ ✓
gemini-3.1-fast A fast edition of Gemini, a cost-effective multimodal model ✓
gemini-3.1-thinking A thinking edition of Gemini, with a built-in reasoning chain, suited to complex problem solving ✓ ✓

Gemini 2.5 series

Model ID Description Vision Extended thinking
gemini-2.5-pro-official Gemini 2.5 Pro, powerful reasoning and long-context capabilities ✓ ✓
gemini-2.5-flash-official Gemini 2.5 Flash, high-speed reasoning with thinking mode support ✓ ✓
gemini-2.5-flash-lite-official Gemini 2.5 Flash Lite, ultra-low cost with ultra-fast responses ✓

Gemini 2.0 series

Model ID Description Vision Extended thinking
gemini-2.0-flash-official Gemini 2.0 Flash, the classic fast model ✓
gemini-2.0-flash-lite-official Gemini 2.0 Flash Lite, the lowest-cost lightweight model ✓

DeepSeek

Model ID Description Vision Extended thinking
deepseek-v3.2 DeepSeek's flagship model, supporting extended thinking, excelling at code and complex reasoning ✓

Alibaba / Qwen

Model ID Description Vision Extended thinking
qwen3-max The Qwen flagship model, supporting extended thinking, excelling at complex reasoning and code generation ✓
qwen3.5-plus An enhanced Qwen edition, with a balanced trade-off of performance and cost
qwen3.5-flash An ultra-fast Qwen edition, with ultra-low latency, the top choice for high-throughput scenarios

Zhipu / GLM

Model ID Description Vision Extended thinking
glm-5 Zhipu's flagship model, with excellent Chinese understanding and generation

Moonshot / Kimi

Model ID Description Vision Extended thinking
kimi-k2.5 Moonshot's flagship model, with ultra-long context and outstanding Chinese capabilities

MiniMax

Model ID Description Vision Extended thinking
MiniMax-M2.5 MiniMax's flagship model, with exceptional value