seed-tts-2.0 Text-to-Speech

Convert text into speech for narration, voiceovers, and voice interactions. This guide uses MaiToken's native HTTP unidirectional streaming endpoint: submit the text in one request and receive audio data as a stream.

POST/api/v3/tts/unidirectional

Prerequisites

  1. Create an API Key in the MaiToken Console.
  2. Make sure the groups selected for your API Key allow access to seed-tts-2.0 and that your account has sufficient available balance.
  3. Prepare your text and a voice ID supported by the model.

Replace the placeholder API Key in the examples with your own production API Key. An admin login token cannot be used as a model API Key.

Endpoint

POST https://{BASE_URL}/api/v3/tts/unidirectional

Use this complete native endpoint path. Do not prepend /v1 to the path.

Request Headers

Header Required Description
X-Api-Key Yes Your MaiToken API Key, without the Bearer prefix
X-Api-Resource-Id Yes Set to seed-tts-2.0
Content-Type Yes application/json; charset=utf-8

Request Body

The sample text below means “Hello, welcome to MaiToken text-to-speech.” It is kept in Chinese to preserve the original voice example.

{  "req_params": {    "text": "你好,欢迎使用 MaiToken 语音合成。",    "speaker": "zh_female_vv_uranus_bigtts",    "audio_params": {      "format": "mp3",      "sample_rate": 24000    }  }}
Parameter Type Description
req_params.text string Text to synthesize. Must contain speakable content and use UTF-8 encoding.
req_params.speaker string Voice ID. Must match the model version. The example uses zh_female_vv_uranus_bigtts.
req_params.audio_params.format string The example explicitly sets this to mp3.
req_params.audio_params.sample_rate integer The example explicitly sets the sample rate to 24000 Hz.

This endpoint selects the model through the request header. You do not need an additional top-level model field in the request body. Voice availability depends on the permissions of the connected service resource.

Bash / cURL Example

This example works in Bash and Git Bash. The test text uses JSON Unicode escapes for “你好” (“Hello”) to avoid terminal encoding issues.

export MAITOKEN_API_KEY='YOUR_PRODUCTION_API_KEY' curl --fail-with-body --silent --show-error --no-buffer \  'https://{BASE_URL}/api/v3/tts/unidirectional' \  -H "X-Api-Key: ${MAITOKEN_API_KEY}" \  -H 'X-Api-Resource-Id: seed-tts-2.0' \  -H 'Content-Type: application/json; charset=utf-8' \  --data-binary '{"req_params":{"text":"\u4f60\u597d","speaker":"zh_female_vv_uranus_bigtts","audio_params":{"format":"mp3","sample_rate":24000}}}' \  --output tts-response.jsonstream

Do not put spaces after the trailing backslash \ on each line. The response file contains a JSON stream; renaming it to .mp3 will not make it playable.

Response Format and Audio Decoding

The HTTP response body contains multiple JSON objects. The data field in each audio object contains Base64-encoded audio. Decode these fragments and concatenate them in order. JSON objects may be separated by newlines or appear directly next to each other. Do not assume that one network chunk contains exactly one complete JSON object.

The following illustrates the response structure. <Base64 audio fragment> is a placeholder, not playable data:

{"code":0,"message":"","data":"<Base64 audio fragment>"}{"code":0,"message":"","data":"<Base64 audio fragment>"}{"code":20000000,"message":"ok","data":null,"usage":{"text_words":2}}

code=0 indicates a normal successful data frame. code=20000000 indicates successful completion of synthesis. usage.text_words contains usage reported by the upstream service and may appear only at the end. Check application-level error codes even when the HTTP status is 200. See the upstream HTTP unidirectional streaming documentation for the response protocol.

Complete Python Example: Generate and Save an MP3

This example uses only the Python standard library; no SDK installation is required. It receives the complete response, parses the consecutive JSON objects, checks the application-level status, and then writes the audio file. It is intended for integration verification, not real-time playback.

Save the following as seed_tts.py:

import base64import jsonimport osfrom pathlib import Pathfrom urllib.error import HTTPErrorfrom urllib.request import Request, urlopen  def decode_audio(raw):    decoder = json.JSONDecoder()    offset = 0    audio = bytearray()    usage = None    finished = False     while offset < len(raw):        if raw[offset].isspace():            offset += 1            continue        frame, offset = decoder.raw_decode(raw, offset)        if not isinstance(frame, dict):            raise RuntimeError("Response frame is not a JSON object")        code = frame.get("code")        if code not in (0, 20000000):            raise RuntimeError(f"Synthesis failed: {code}, {frame.get('message', frame)}")        if frame.get("data"):            audio.extend(base64.b64decode(frame["data"], validate=True))        if isinstance(frame.get("usage"), dict):            usage = frame["usage"]        if code == 20000000:            finished = True     if not finished:        raise RuntimeError("No successful completion frame received; check for a connection interruption")    if not audio:        raise RuntimeError("The response contains no audio data")    return bytes(audio), usage  def main():    payload = {        "req_params": {            "text": "你好,欢迎使用 MaiToken 语音合成。",            "speaker": "zh_female_vv_uranus_bigtts",            "audio_params": {"format": "mp3", "sample_rate": 24000},        }    }    request = Request(        "https://{BASE_URL}/api/v3/tts/unidirectional",        data=json.dumps(payload, ensure_ascii=False).encode("utf-8"),        headers={            "X-Api-Key": os.environ["MAITOKEN_API_KEY"],            "X-Api-Resource-Id": "seed-tts-2.0",            "Content-Type": "application/json; charset=utf-8",        },        method="POST",    )    try:        with urlopen(request, timeout=120) as response:            raw = response.read().decode("utf-8")    except HTTPError as exc:        detail = exc.read().decode("utf-8", errors="replace")        raise RuntimeError(f"HTTP {exc.code}: {detail}") from exc     audio, usage = decode_audio(raw)    Path("speech.mp3").write_bytes(audio)    print(f"Saved speech.mp3: {len(audio)} bytes; usage: {usage}")  if __name__ == "__main__":    main()

Run it with:

export MAITOKEN_API_KEY='YOUR_PRODUCTION_API_KEY'python seed_tts.py

If you already saved the response using the cURL example, reuse the decoder to reconstruct the audio without making another API request:

python -c 'from pathlib import Path; from seed_tts import decode_audio; audio, usage = decode_audio(Path("tts-response.jsonstream").read_text(encoding="utf-8")); Path("speech.mp3").write_bytes(audio); print(usage)'

Usage and Billing

This model is billed by character usage, with prices expressed per 10,000 characters. The applicable selling price depends on MaiToken's current model pricing and your account's group.

Charge = Billable characters / 10,000 × Selling price per 10,000 characters × Group multiplier × User multiplier

For example, if the price is 0.08 credits per 10,000 characters, the billable usage is 1,000 characters, and both multipliers are 1, the charge is 0.008 credits. This price is illustrative and is not a production price quote.

An estimated amount may be reserved when a request starts. The final charge is settled according to usage, and any difference is reconciled. Review the reserved amount separately from the final charge. Use the settlement record in the call logs as the authoritative usage record; do not calculate character usage from UTF-8 byte length.

Troubleshooting

Issue What to check
curl: (3) URL rejected Check URL quotation marks, invisible characters, and line-continuation backslashes. If necessary, put the command on a single line.
unsupported_endpoint Use the complete path /api/v3/tts/unidirectional. Do not omit /api/v3 or add /v1.
No readable text! Make sure req_params.text is nonempty and contains speakable content. Use the Unicode-escaped test text to rule out encoding problems.
HTTP 401 Use a valid MaiToken API Key, not an admin login token.
Insufficient balance Check available balance and funds reserved by ongoing requests.
Voice permission error Verify that speaker identifies a voice supported by the model and authorized for the service resource.
HTTP 429 / concurrency limit Reduce concurrency and use retry backoff according to the response details.
Audio file cannot be played Base64-decode each data fragment and concatenate the decoded bytes in order. Do not save the JSON response directly as an MP3.
Interrupted connection / missing completion frame Do not treat partial audio as a complete success. Check the call status and billing record before deciding whether to submit another request.

Sources and Verification

  • MaiToken API Overview: production API domain and integration entry point.
  • MaiToken Speech Synthesis Page: the page body was empty when checked; this guide supplies the missing usage documentation.
  • Upstream HTTP Unidirectional Streaming Protocol: audio frame and completion frame semantics.
  • Platform endpoint paths, request forwarding, character usage, and billing units were checked against the current gateway and metering implementations. No paid synthesis request was made using a real API Key to prepare this guide.

Verified on: September 24, 2026.