MiniMax H3 Video Generation API

This API supports the official minimax-h3 model and three self-hosted models: minimax-h3-base, minimax-h3-base-fast, and minimax-h3-mini.

All models share the same submission, status, and download endpoints. For self-hosted models, the platform automatically translates unified parameters into upstream-native fields. You can continue using duration, image, and reference_image_urls without switching to native fields such as seconds or images.

1. Endpoints

Operation Method Path
Submit a generation task POST /v1/video/generations
Submit a generation task (equivalent endpoint) POST /v1/videos
Query task status GET /v1/video/generations/{task_id}
Download the generated video GET /v1/videos/{task_id}/content

Video generation is asynchronous:

  1. Submit a generation request.
  2. Retrieve the task_id from the response.
  3. Query the task until generation is complete.
  4. Download the generated video.

{BASE_URL} represents your platform’s API base URL. Authentication requirements, response schemas, task status values, and error codes are not included in the supplied specification. Refer to your platform’s documentation for these details.

2. Models and Capabilities

Capability MiniMax-H3 (Official) Self-Hosted Models
Model IDs MiniMax-H3 minimax-h3-base, minimax-h3-base-fast, minimax-h3-mini
Duration 4–15 seconds; default: 5 seconds 5–15 seconds
Resolution 768P (default), 2K 720p only; 768P is converted to 720p
First and last frames Supported Supported
Reference images Up to 9 images Up to 9 images; maximum 20 MB per image
Reference videos Up to 3 clips Not supported
Reference audio Up to 3 clips Up to 3 clips; maximum 50 MB per clip
Output details Includes native stereo audio MP4, 24 fps

The three self-hosted models use the same parameters and support the same capabilities. They differ in generation speed and pricing.

All reference media URLs must be publicly accessible from mainland China. Overseas direct links that cannot be accessed from mainland China are not supported.

3. Submit a Generation Task

POST {BASE_URL}/v1/video/generationsContent-Type: application/json

Equivalent endpoint:

POST {BASE_URL}/v1/videos

3.1 Unified Parameters

Parameter Type Description
model string Required. The model ID.
prompt string Describes the video to generate. Required for self-hosted models, with a maximum length of 30,000 characters.
duration number Video duration in seconds. Official: 4–15. Self-hosted: 5–15.
resolution string Official: 768P or 2K. For self-hosted models, explicitly use 720p.
aspect_ratio string Output aspect ratio. See the rules below.
image string URL of a first-frame image or a single reference image.
images array of strings Image URLs; supported as a compatibility field.
last_frame_image_url string URL of the last-frame image. Use with image for first-and-last-frame generation.
reference_image_urls array of strings Reference image URLs. Up to 9 images.
reference_video_urls array of strings Reference video URLs. Official model only; up to 3 clips.
reference_audio_urls array of strings Reference audio URLs. Up to 3 clips.
generation_mode string Optional. t2v, i2v, or r2v. If omitted, the platform determines the mode from the supplied media.

Explicitly set duration and resolution to avoid relying on model-specific defaults.

3.2 Aspect Ratio

Official model

  • Text-to-video defaults to 16:9.
  • adaptive is recommended for first-and-last-frame or multi-reference generation.

Self-hosted models

Supported fixed aspect ratios:

16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 4:5, 5:4

When aspect_ratio is set to adaptive, the platform omits the aspect-ratio field from the upstream request, allowing the upstream service to determine the output ratio.

3.3 Self-Hosted Options and Limits

Parameter or Media Rule
prompt_optimization Optional boolean; defaults to false. When true, the upstream service rewrites the prompt using the official H3 prompt structure.
Reference images Up to 9 images, no larger than 20 MB each.
Reference audio Up to 3 clips, no larger than 50 MB each.
Multi-reference generation Requires at least one reference image or one reference audio clip.
First and last frames Use image and last_frame_image_url for the first and last frames, respectively.
Reference videos Not supported. Do not send reference_video_urls.

The platform does not reject out-of-range parameters for self-hosted models. Unsupported values such as a 4-second duration, 2K, or reference videos are forwarded upstream, where they may be ignored or cause an error. Billing is still provisionally deducted based on the duration declared in your request.

For self-hosted models, use 5–15 seconds, set resolution to 720p, and omit reference videos.

4. Request Examples

The following request bodies work with either submission endpoint. Replace the example media URLs with real URLs accessible from mainland China.

4.1 Official Model: Reference Image Generation

{  "model": "minimax-h3",  "prompt": "The character from reference image 1 runs through a rainy street at night, with cinematic camera work.",  "resolution": "768P",  "duration": 8,  "aspect_ratio": "16:9",  "reference_image_urls": [    "https://example.com/character.png"  ]}

4.2 Self-Hosted: Text-to-Video

{  "model": "minimax-h3-base-fast",  "prompt": "At dawn by the sea, a seagull glides over the water as the camera slowly moves forward in soft natural light.",  "resolution": "720p",  "duration": 8,  "aspect_ratio": "16:9",  "generation_mode": "t2v"}

4.3 Self-Hosted: First-Frame Generation

{  "model": "minimax-h3-base",  "prompt": "The character in the image runs through a rainy street at night, with cinematic camera work.",  "resolution": "720p",  "aspect_ratio": "16:9",  "duration": 8,  "image": "https://example.com/character.png"}

4.4 Self-Hosted: First-and-Last-Frame Generation

{  "model": "minimax-h3-base",  "prompt": "The character walks naturally from one end of the street to the other. Maintain a consistent appearance and use a smooth tracking shot.",  "resolution": "720p",  "duration": 8,  "aspect_ratio": "adaptive",  "image": "https://example.com/first-frame.png",  "last_frame_image_url": "https://example.com/last-frame.png"}

The platform passes the first and last frames to the upstream service in that order and automatically sets the upstream mode to fl2va.

4.5 Self-Hosted: Image and Audio References

{  "model": "minimax-h3-mini",  "prompt": "Preserve the character's appearance from image 1 and make the character speak with the tone and rhythm of audio 1.",  "resolution": "720p",  "duration": 10,  "reference_image_urls": [    "https://example.com/character.png"  ],  "reference_audio_urls": [    "https://example.com/voice.mp3"  ]}

4.6 Submit with cURL

Save a request body as request.json, then run:

curl --request POST "${BASE_URL}/v1/video/generations" \  --header "Content-Type: application/json" \  --data-binary @request.json

Add the authentication headers required by your platform.

5. Query Status and Download the Video

5.1 Query Task Status

Use the task ID returned by the submission request:

GET {BASE_URL}/v1/video/generations/{task_id}
curl "${BASE_URL}/v1/video/generations/${TASK_ID}"

Check the response to determine whether generation has completed. The exact status fields, status values, and error details depend on the platform’s response schema.

5.2 Download the Result

Once generation is complete:

GET {BASE_URL}/v1/videos/{task_id}/content
curl --location --fail \  "${BASE_URL}/v1/videos/${TASK_ID}/content" \  --output video.mp4

Include the platform’s required authentication information in both status and download requests.

6. Self-Hosted Parameter Translation

The platform performs these conversions automatically.

Meaning Unified Input Upstream Field or Behavior
Duration duration: 8 "seconds": "8"
Reference images or first frame reference_image_urls, image, images, input_reference Merged in the listed order and deduplicated into images.
Last frame last_frame_image_url Appended to images; automatically sets mode=fl2va.
Reference audio reference_audio_urls audios
Generation mode generation_mode Translated to mode; automatically inferred from supplied media if omitted.
Resolution resolution: "768P" resolution: "720p"; no equivalent tier exists for 2K.
Adaptive aspect ratio aspect_ratio: "adaptive" The aspect-ratio field is omitted.
Aspect-ratio alias ratio Normalized to aspect_ratio.
Reference videos reference_video_urls Unsupported; forwarded unchanged for upstream handling.

For first-and-last-frame generation, use image and last_frame_image_url to specify the two frames. Avoid mixing additional image fields into the request, as merging them may affect image order.

7. Native Upstream Field Compatibility

Self-hosted models also accept native upstream fields directly. The platform leaves these native fields unchanged.

Native Field Accepted Format or Values
seconds String or number, such as "8" or 8.
images Array of image URLs.
audios Array of audio URLs.
mode t2va, i2va, fl2va, l2va, or ref2va.

Example:

{  "model": "minimax-h3-base",  "prompt": "The character in the image runs through a rainy street at night, with cinematic camera work.",  "resolution": "720p",  "aspect_ratio": "16:9",  "seconds": "8",  "images": [    "https://example.com/character.png"  ],  "mode": "i2va"}

For new integrations, use the unified parameters. Conflict precedence between equivalent unified and native fields is not specified, so avoid sending both forms of the same setting—for example, duration and seconds—in one request.