The platform offers two ways to call the API, switching models via the model parameter:
| API format |
Endpoint |
Notes |
| Chat Completions (default) |
POST /v1/chat/completions |
Classic conversation endpoint, compatible with all OpenAI SDKs. View docs |
| Responses (new) |
POST /v1/responses |
OpenAI's next-generation Agentic API, supporting function calling, built-in tools, and server-side multi-turn context. View docs |
Tip: Models marked Responses Only support only the Responses API, not Chat Completions. All other models are called through Chat Completions by default.
Anthropic / Claude
| Model ID |
Description |
Vision |
Extended thinking |
claude-haiku-4-5 |
The fastest, lowest-cost Claude model, ideal for high-concurrency lightweight tasks |
✓ |
|
claude-sonnet-4-6 |
The best balance of performance and cost, suitable for most business scenarios |
✓ |
|
claude-opus-4-6 |
Claude's flagship model, suitable for complex reasoning and high-quality creation |
✓ |
|
OpenAI / GPT-5
GPT-5 Official (direct official connection)
Connected directly through the Azure OpenAI official API, offering the highest stability and compliance.
| Model ID |
Description |
Vision |
Extended thinking |
API format |
gpt-5-pro-official |
GPT-5 Pro direct official connection, OpenAI's most powerful flagship reasoning model |
✓ |
✓ |
Responses Only |
gpt-5.2-official |
GPT-5.2 direct official connection, multimodal enhanced edition |
✓ |
✓ |
All |
gpt-5.3-codex-official |
GPT-5.3 Codex direct official connection, the latest code-optimized edition, supporting function calling and built-in tools |
✓ |
✓ |
Responses Only |
GPT-5 Standard
| Model ID |
Description |
Vision |
Extended thinking |
gpt-5 |
OpenAI's flagship model, powerful general reasoning and multimodal capabilities |
✓ |
|
gpt-5-codex |
A GPT-5 optimized for code, excelling at code generation and debugging |
|
|
gpt-5-codex-mini |
A lightweight code model with ultra-fast responses at low cost |
|
|
gpt-5.1 |
An enhanced GPT-5 with improved instruction following and reasoning |
✓ |
|
gpt-5.1-codex |
The code-specialized edition of GPT-5.1 |
|
|
gpt-5.1-codex-max |
The flagship code edition of GPT-5.1, with the strongest coding capabilities |
|
|
gpt-5.1-codex-mini |
A lightweight code edition of GPT-5.1, fast and low-cost |
|
|
gpt-5.2 |
The second generation of GPT-5, multimodal enhanced |
✓ |
|
gpt-5.2-codex |
The code-specialized edition of GPT-5.2 |
|
|
gpt-5.3-codex |
The latest code edition of GPT-5, with continuously improved coding capabilities |
|
|
gpt-5.3-codex-spark |
The GPT-5.3 Codex Spark variant, focused on high-speed code generation |
|
|
gpt-5.4 |
The fourth generation of GPT-5, with stronger general reasoning and multimodal capabilities |
✓ |
|
gpt-5.5 |
GPT-5.5, with stronger general reasoning, coding, and multimodal capabilities |
✓ |
|
gpt-5.6-sol |
The GPT-5.6 flagship model, suitable for complex reasoning, high-value code, and long-context tasks |
✓ |
✓ |
gpt-5.6-terra |
The GPT-5.6 balanced model, suitable for general production traffic and agent workflows |
✓ |
✓ |
gpt-5.6-luna |
The GPT-5.6 high-throughput model, suitable for classification, extraction, routing, and batch processing |
✓ |
✓ |
Google / Gemini
Gemini 3 series
| Model ID |
Description |
Vision |
Extended thinking |
gemini-3.1-pro-preview-official |
Gemini 3.1 Pro, the latest-generation flagship reasoning model |
✓ |
✓ |
gemini-3.1-flash-lite-preview-official |
Gemini 3.1 Flash Lite, the most cost-effective lightweight model, with ultra-fast responses |
✓ |
|
gemini-3-pro-official |
Gemini 3 Pro flagship, the strongest reasoning and multimodal capabilities |
✓ |
✓ |
gemini-3-pro-preview-official |
The Gemini 3 Pro Preview edition (same as gemini-3-pro-official) |
✓ |
✓ |
gemini-3-flash-official |
Gemini 3 Flash, fast and efficient, a great value choice |
✓ |
✓ |
gemini-3-flash-preview-official |
The Gemini 3 Flash Preview edition (same as gemini-3-flash-official) |
✓ |
✓ |
gemini-3.1-fast |
A fast edition of Gemini, a cost-effective multimodal model |
✓ |
|
gemini-3.1-thinking |
A thinking edition of Gemini, with a built-in reasoning chain, suited to complex problem solving |
✓ |
✓ |
Gemini 2.5 series
| Model ID |
Description |
Vision |
Extended thinking |
gemini-2.5-pro-official |
Gemini 2.5 Pro, powerful reasoning and long-context capabilities |
✓ |
✓ |
gemini-2.5-flash-official |
Gemini 2.5 Flash, high-speed reasoning with thinking mode support |
✓ |
✓ |
gemini-2.5-flash-lite-official |
Gemini 2.5 Flash Lite, ultra-low cost with ultra-fast responses |
✓ |
|
Gemini 2.0 series
| Model ID |
Description |
Vision |
Extended thinking |
gemini-2.0-flash-official |
Gemini 2.0 Flash, the classic fast model |
✓ |
|
gemini-2.0-flash-lite-official |
Gemini 2.0 Flash Lite, the lowest-cost lightweight model |
✓ |
|
DeepSeek
| Model ID |
Description |
Vision |
Extended thinking |
deepseek-v3.2 |
DeepSeek's flagship model, supporting extended thinking, excelling at code and complex reasoning |
|
✓ |
Alibaba / Qwen
| Model ID |
Description |
Vision |
Extended thinking |
qwen3-max |
The Qwen flagship model, supporting extended thinking, excelling at complex reasoning and code generation |
|
✓ |
qwen3.5-plus |
An enhanced Qwen edition, with a balanced trade-off of performance and cost |
|
|
qwen3.5-flash |
An ultra-fast Qwen edition, with ultra-low latency, the top choice for high-throughput scenarios |
|
|
Zhipu / GLM
| Model ID |
Description |
Vision |
Extended thinking |
glm-5 |
Zhipu's flagship model, with excellent Chinese understanding and generation |
|
|
Moonshot / Kimi
| Model ID |
Description |
Vision |
Extended thinking |
kimi-k2.5 |
Moonshot's flagship model, with ultra-long context and outstanding Chinese capabilities |
|
|
MiniMax
| Model ID |
Description |
Vision |
Extended thinking |
MiniMax-M2.5 |
MiniMax's flagship model, with exceptional value |
|
|