Advanced functions

Q: Does it support RAG (Retrieval-Augmented Generation)?

A: The platform provides the following solution:

Self-built RAG: Use our Embedding API to vectorize documents → store them in your vector database (such as Milvus/Pinecone) → retrieve relevant passages → add them to the prompt and call the chat model

Q: Does it support JSON Mode / structured output?

A: Yes. Set response_format={"type": "json_object"} to make the model output valid JSON:

response = client.chat.completions.create(    model="[Model Name]",    messages=[        {"role": "system", "content": "Please output in JSON format, including the name and age fields"},        {"role": "user", "content": "Zhang San, 25 years old"}    ],    response_format={"type": "json_object"})

Note: The system prompt must include a "JSON"-related instruction; otherwise, an error may occur.

Q: What is the specific format for Vision (image understanding)?

A: Two methods are supported for providing images:

Method 1: Use an image URL
image_by_url = {    "type": "image_url",    "image_url": {        "url": "https://example.com/img.jpg"    }}
Method 2: Use Base64
image_by_base64 = {    "type": "image_url",    "image_url": {        "url": "data:image/jpeg;base64,/9j/4AAQ..."    }}

Supported formats: JPEG / PNG / GIF / WebP

  • Maximum size per image: 4 MB

  • Maximum of 1 image per request

  • Image token consumption is related to resolution; higher-resolution images consume more tokens