Calling parameters

A: The core parameters are explained below.

Q: What is the difference between temperature and top_p, and how should I configure them?

A:

  • temperature: Controls the overall randomness. Lower values produce more deterministic and conservative outputs, while higher values produce more diverse and creative outputs.

  • top_p: Controls nucleus sampling. The model samples only from candidate tokens whose cumulative probability reaches the specified top_p value.

  • Generally, adjust only one of these parameters and leave the other at its default value of 1.0.

  • Recommended configurations:

    • Customer service / FAQ: temperature=0.1, top_p=1.0 for stable and accurate responses

    • Content creation: temperature=0.8, top_p=0.95 for richer and more diverse responses

    • Code generation: temperature=0.2, top_p=1.0 for precise and reliable output

Q: What is a system prompt, and how do I write an effective one?

A: A system prompt is a global instruction placed at the beginning of a conversation. It defines the model's role, behavioral rules, and output format.

Customer service example:

You are the intelligent customer support assistant for [Company Name]. Follow these rules:Use a friendly and professional tone, and address customers politely.Keep responses clear and concise, using no more than 200 words.For questions you cannot answer, direct the customer to human support at 400-XXX-XXXX.Do not answer questions unrelated to [Company/Product].Do not fabricate information. If you are unsure, state that clearly.

Optimization recommendations:

  • Keep the system prompt within 500 Tokens. Longer prompts increase the input Token cost of every request.

  • Clearly define the role, boundaries, and formatting requirements.

  • Use a numbered list for rules instead of long paragraphs.

  • Add two or three few-shot examples to guide the desired output style.

Q: What is enable_thinking (deep thinking mode)?

A: When enabled, the model performs internal reasoning before producing the final answer. It is suitable for complex logic, mathematics, coding, and similar tasks.

  • To enable it, add "enable_thinking": true to the request parameters.

  • The thinking content is returned in the reasoning_content field.

  • ⚠️ Note: The thinking process consumes additional Tokens and increases response latency.

  • Deep thinking mode is not recommended for basic conversations or customer service scenarios because it increases both cost and latency.

Q: How do I implement multi-turn conversations?

A: Pass the conversation history to the messages array in chronological order:

messages = [    {"role": "system", "content": "You are a customer support assistant."},    {"role": "user", "content": "How do I request a refund?"},    {"role": "assistant", "content": "The refund process is as follows..."},    {"role": "user", "content": "How long will it take to receive the refund?"}  # Current question]

💡 Cost control recommendation: When the conversation history becomes too long, truncate it by keeping only the most recent N rounds, or compress earlier messages into a summary to prevent input Token usage from growing continuously.