Q: What do the parameters mean, and what values are recommended?
A: The core parameters are explained below.
Q: What is the difference between temperature and top_p, and how should I configure them?
A:
temperature: Controls the overall randomness. Lower values produce more deterministic and conservative outputs, while higher values produce more diverse and creative outputs.top_p: Controls nucleus sampling. The model samples only from candidate tokens whose cumulative probability reaches the specifiedtop_pvalue.Generally, adjust only one of these parameters and leave the other at its default value of
1.0.Recommended configurations:
Customer service / FAQ:
temperature=0.1, top_p=1.0for stable and accurate responsesContent creation:
temperature=0.8, top_p=0.95for richer and more diverse responsesCode generation:
temperature=0.2, top_p=1.0for precise and reliable output
Q: What is a system prompt, and how do I write an effective one?
A: A system prompt is a global instruction placed at the beginning of a conversation. It defines the model's role, behavioral rules, and output format.
Customer service example:
You are the intelligent customer support assistant for [Company Name]. Follow these rules:Use a friendly and professional tone, and address customers politely.Keep responses clear and concise, using no more than 200 words.For questions you cannot answer, direct the customer to human support at 400-XXX-XXXX.Do not answer questions unrelated to [Company/Product].Do not fabricate information. If you are unsure, state that clearly.Optimization recommendations:
Keep the system prompt within 500 Tokens. Longer prompts increase the input Token cost of every request.
Clearly define the role, boundaries, and formatting requirements.
Use a numbered list for rules instead of long paragraphs.
Add two or three few-shot examples to guide the desired output style.
Q: What is enable_thinking (deep thinking mode)?
A: When enabled, the model performs internal reasoning before producing the final answer. It is suitable for complex logic, mathematics, coding, and similar tasks.
To enable it, add
"enable_thinking": trueto the request parameters.The thinking content is returned in the
reasoning_contentfield.⚠️ Note: The thinking process consumes additional Tokens and increases response latency.
Deep thinking mode is not recommended for basic conversations or customer service scenarios because it increases both cost and latency.
Q: How do I implement multi-turn conversations?
A: Pass the conversation history to the messages array in chronological order:
messages = [ {"role": "system", "content": "You are a customer support assistant."}, {"role": "user", "content": "How do I request a refund?"}, {"role": "assistant", "content": "The refund process is as follows..."}, {"role": "user", "content": "How long will it take to receive the refund?"} # Current question]💡 Cost control recommendation: When the conversation history becomes too long, truncate it by keeping only the most recent N rounds, or compress earlier messages into a summary to prevent input Token usage from growing continuously.
