Prerequisites
- Create an API key in the MaiToken Console.
- Make sure the group selected for the key allows
qwen-audio-3.0-tts-plusand has sufficient available balance. - Select a voice supported by the model. This guide uses
longanlingxin, which was successfully used in the production HTTP test described below. For other voices, see the official voice list.
Examples use environment variables or placeholder keys only. An admin login token cannot replace a model API key. Chinese synthesis inputs are retained from the original guide so the documented test results remain applicable.
Endpoints and Invocation Methods
| Method | Production endpoint | Input and output |
|---|---|---|
| Bidirectional WebSocket streaming | wss://api.maitoken.com/api-ws/v1/inference |
Submit text segments and receive JSON events and binary audio frames |
| Non-realtime HTTP | https://api.maitoken.com/api/v1/services/audio/tts/SpeechSynthesizer |
Submit text in one request and receive JSON containing an audio URL |
Use the complete native path without adding another /v1. Do not pass an HTTP URL directly to a WebSocket client. The WebSocket route for this model is /api-ws/v1/inference; the model name belongs in the first run-task message.
/api-ws/v1/realtime?model=... uses a different Qwen-TTS Realtime session protocol with events such as session.update. Do not mix it with the run-task protocol in this guide. Availability of other models depends on the model list and your key's permissions.
Authentication
Include the following header in the WebSocket HTTP handshake or HTTP synthesis request:
Authorization: Bearer YOUR_MAITOKEN_API_KEYWebSocket requests do not require X-Api-Resource-Id, X-DashScope-Async, or X-DashScope-SSE. Send JSON control messages as WebSocket text frames and receive audio as binary frames.
The browser's native WebSocket constructor cannot set a custom Authorization header. Web applications should connect through their own backend. Do not put the API key in the URL as a substitute for the header. The Python example below sets the handshake header directly.
WebSocket Workflow
- Establish a connection with the API key in the handshake header.
- Generate a new task UUID and send
run-task. - Wait for the server's
task-startedevent before sending text. - Submit text sequentially using one or more
continue-taskmessages. - Send
finish-taskwhen all text has been submitted, then continue receiving the remaining audio. - Treat the task as successfully completed only after receiving
task-finished. Handletask-failedas a failure.
Use the same task_id for every message in a task. MaiToken currently allows only one synthesis task per inference connection. Establish a new connection for the next task; do not apply the upstream documentation's connection reuse behavior to this gateway.
1. Start a Task: run-task
{ "header": { "action": "run-task", "task_id": "dbe53a5f-257f-4302-a9c1-1cf1d4790b52", "streaming": "duplex" }, "payload": { "task_group": "audio", "task": "tts", "function": "SpeechSynthesizer", "model": "qwen-audio-3.0-tts-plus", "parameters": { "voice": "longanlingxin", "format": "mp3", "sample_rate": 24000 }, "input": {} }}| Parameter | Type | Description |
|---|---|---|
header.action |
string | Set torun-task when starting a task |
header.task_id |
string | Generate a UUID for each new task and reuse it in subsequent messages |
header.streaming |
string | Set toduplex for bidirectional streaming |
payload.task_group |
string | Set toaudio |
payload.task |
string | Set totts |
payload.function |
string | Set toSpeechSynthesizer |
payload.model |
string | Exact model ID; this guide usesqwen-audio-3.0-tts-plus |
payload.parameters.voice |
string | A supported voice ID accessible with the current model |
payload.parameters.format |
string | This guide usesmp3; adjust audio handling if you choose another format |
payload.parameters.sample_rate |
integer | This guide uses24000 Hz; refer to the model documentation for supported values |
payload.input |
object | Pass{} when starting a bidirectional task; submit text in subsequent messages |
The model and voice are fixed when the task starts and cannot be changed in subsequent continue-task messages. Voice access depends on the upstream resources in use. Custom voices should first be created and registered through the corresponding MaiToken voice management API.
2. Submit Text: continue-task
After receiving task-started, send:
{ "header": { "action": "continue-task", "task_id": "dbe53a5f-257f-4302-a9c1-1cf1d4790b52", "streaming": "duplex" }, "payload": { "input": { "text": "你好,欢迎使用 MaiToken。" } }}Submit additional text using messages with the same structure. Short, natural sentences are recommended. Production applications should receive audio concurrently to avoid a buffer backlog caused by sending without reading. Do not assume that each text submission produces exactly one audio frame.
3. Finish Input: finish-task
{ "header": { "action": "finish-task", "task_id": "dbe53a5f-257f-4302-a9c1-1cf1d4790b52", "streaming": "duplex" }, "payload": { "input": {} }}finish-task means that all text has been sent, not that all audio has been received. Do not disconnect immediately after sending it.
Responses and Audio Reconstruction
| Response | Meaning and handling |
|---|---|
JSON:header.event=task-started |
The task has started; text can now be submitted |
| Binary frame | Audio chunk; write it to a file or pass it to a decoder in receive order |
JSON:header.event=result-generated |
An intermediate result or usage event; inspectpayload as needed |
JSON:header.event=task-finished |
Synthesis completed successfully; close the connection and finalize the file |
JSON:header.event=task-failed |
Inspectheader.error_code and header.error_message |
Illustrative completion event:
{ "header": { "task_id": "dbe53a5f-257f-4302-a9c1-1cf1d4790b52", "event": "task-finished" }, "payload": { "usage": { "characters": 26 } }}This illustrates the event structure only. Actual fields and the events carrying usage depend on the upstream response. payload.usage.characters reports character usage. A task may report cumulative values more than once; do not add those cumulative values together.
With the MP3 format used here, concatenate binary frames in order. These frames are not Base64 and do not need Base64 decoding. Never write JSON control frames into the audio file. If you switch to PCM, configure the player for the returned sample rate, channel count, and sample format. Renaming a PCM file does not convert it to MP3 or WAV.
Complete Python Example: Receive Streaming Audio and Save MP3
Install the dependency:
python -m pip install websocket-clientSave the following as qwen_tts_stream.py. It receives and saves audio over WebSocket but does not implement audio-device playback. On failure, the .part file is retained for inspection. The final MP3 is created only after the successful completion event is received.
import jsonimport osimport timeimport uuidfrom pathlib import Path import websocket def main(): key = os.environ.get("MAITOKEN_API_KEY", "").strip() if not key: raise RuntimeError("Set MAITOKEN_API_KEY first") task_id = str(uuid.uuid4()) target = Path(f"qwen-{task_id}.mp3") partial = target.with_suffix(".mp3.part") texts = ["你好,欢迎使用 MaiToken。", "这是千问实时语音合成示例。"] started = False finished = False audio_bytes = 0 latest_usage = None ws = websocket.create_connection( "wss://api.maitoken.com/api-ws/v1/inference", header={"Authorization": f"Bearer {key}"}, timeout=30, ) def send(action, payload): ws.send(json.dumps({ "header": { "action": action, "task_id": task_id, "streaming": "duplex", }, "payload": payload, }, ensure_ascii=False)) try: print("Task ID:", task_id) send("run-task", { "task_group": "audio", "task": "tts", "function": "SpeechSynthesizer", "model": "qwen-audio-3.0-tts-plus", "parameters": { "voice": "longanlingxin", "format": "mp3", "sample_rate": 24000, }, "input": {}, }) deadline = time.monotonic() + 120 with partial.open("xb") as output: while not finished: remaining = deadline - time.monotonic() if remaining <= 0: raise TimeoutError("Task exceeded the example's 120-second total timeout") ws.settimeout(min(30, remaining)) frame = ws.recv() if isinstance(frame, bytes): output.write(frame) audio_bytes += len(frame) continue if not frame: raise RuntimeError("Connection closed before task-finished was received") event = json.loads(frame) header = event.get("header") or {} if header.get("task_id") not in (None, "", task_id): raise RuntimeError("Received an event for a different task") payload = event.get("payload") or {} usage = payload.get("usage") if isinstance(usage, dict): latest_usage = usage name = header.get("event") if name == "task-started" and not started: started = True # These two short sentences can be submitted immediately. # For continuous LLM output, use separate sending logic while # keeping this receive loop running. Send finish-task after all text. for text in texts: send("continue-task", {"input": {"text": text}}) send("finish-task", {"input": {}}) elif name == "task-failed": raise RuntimeError( f"Task failed: {header.get('error_code')}, " f"{header.get('error_message')}" ) elif name == "task-finished": finished = True if not audio_bytes: raise RuntimeError("Task finished without returning audio data") partial.replace(target) print(f"Saved {target}: {audio_bytes} bytes") print("Upstream usage:", latest_usage) finally: ws.close() if __name__ == "__main__": main()Bash / Git Bash:
export MAITOKEN_API_KEY='YOUR_MAITOKEN_API_KEY'python qwen_tts_stream.pyWindows PowerShell:
$env:MAITOKEN_API_KEY = 'YOUR_MAITOKEN_API_KEY'python qwen_tts_stream.pyThe example uses a 30-second timeout per receive operation and a 120-second total task timeout for short-text integration checks. Adjust these values for longer text. A timeout or interruption does not necessarily mean that no charges were incurred; do not retry indefinitely.
Non-Realtime HTTP Invocation
Use the HTTP route if you do not need to submit text incrementally while receiving audio. Note the parameter locations: HTTP places voice, format, and sample_rate inside input, whereas WebSocket places them inside payload.parameters.
Bash / cURL
export MAITOKEN_API_KEY='YOUR_MAITOKEN_API_KEY' curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.maitoken.com/api/v1/services/audio/tts/SpeechSynthesizer' \ -H "Authorization: Bearer ${MAITOKEN_API_KEY}" \ -H 'Content-Type: application/json; charset=utf-8' \ --data-binary '{"model":"qwen-audio-3.0-tts-plus","input":{"text":"你好,这是千问语音合成测试。","voice":"longanlingxin","format":"mp3","sample_rate":24000}}'Troubleshooting
| Symptom | What to check |
|---|---|
| HTTP 401 during handshake | Verify that the MaiToken key is valid and includes theBearer prefix; do not use an admin login token |
Model unavailable ortask-failed |
Check key group permissions, the model ID in the first frame, voice permissions, and the error message; a successful handshake does not mean synthesis succeeded |
404 orunsupported_endpoint |
Check the complete native path, avoid adding/v1 twice, and distinguish inference from realtime |
| No audio after submitting text | Confirm receipt oftask-started, use the same task_id in every frame, and send finish-task after all text |
| Truncated audio | Wait for task-finished instead of disconnecting immediately after sending finish-task |
| MP3 cannot be played | Concatenate only binary audio frames; do not write JSON text or Base64-decode binary frames |
| Changing the model or voice is rejected | Open a new connection and specify them in a new run-task |
| A second run-task is rejected | MaiToken currently allows one task per connection; reconnect for a new task |
| 429 or concurrency limit | Reduce concurrency and use backoff; do not unconditionally replay a whole task after an interruption once audio has been received |
| Character usage differs from text length | Use upstream characters, not byte counts, audio duration, or local string length |
| A custom voice cannot be used | Check that it is registered through MaiToken and that its target model matches the current model |
For SSML, instruction control, and other advanced parameters, follow the Qwen documentation's model compatibility and input constraints. The examples use plain text and do not enable these features by default. AOQ, mentioned in the Qwen documentation, is outside the MaiToken integration methods verified in this guide.
References and Verification Scope
- Qwen realtime speech synthesis: synthesis modes and native WebSocket event flow.
- Current MaiToken speech synthesis documentation: the existing Seed TTS documentation structure and API comparison.
- Qwen-Audio-TTS voice list: voice selection reference.
- MaiToken gateway routing and WebSocket validation implementation: the inference route, model selection from the first frame, and the one-task-per-connection restriction.
Verified on September 29, 2026. The production HTTP call returned status 200, an audio URL, and usage of 26 characters. No additional paid synthesis calls were made while preparing this document. The WebSocket example was written against the official protocol and gateway implementation and passed a static syntax check; it has not been tested in production. HTTP success must not be interpreted as completed WebSocket acceptance testing.
