Q: Does it support RAG (Retrieval-Augmented Generation)?
A: The platform provides the following solution:
Self-built RAG: Use our Embedding API to vectorize documents → store them in your vector database (such as Milvus/Pinecone) → retrieve relevant passages → add them to the prompt and call the chat model
Q: Does it support JSON Mode / structured output?
A: Yes. Set response_format={"type": "json_object"} to make the model output valid JSON:
response = client.chat.completions.create( model="[Model Name]", messages=[ {"role": "system", "content": "Please output in JSON format, including the name and age fields"}, {"role": "user", "content": "Zhang San, 25 years old"} ], response_format={"type": "json_object"})Note: The system prompt must include a "JSON"-related instruction; otherwise, an error may occur.
Q: What is the specific format for Vision (image understanding)?
A: Two methods are supported for providing images:
Method 1: Use an image URL
image_by_url = { "type": "image_url", "image_url": { "url": "https://example.com/img.jpg" }}Method 2: Use Base64
image_by_base64 = { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQ..." }}Supported formats: JPEG / PNG / GIF / WebP
Maximum size per image: 4 MB
Maximum of 1 image per request
Image token consumption is related to resolution; higher-resolution images consume more tokens
