All models share the same submission, status, and download endpoints. For self-hosted models, the platform automatically translates unified parameters into upstream-native fields. You can continue using duration, image, and reference_image_urls without switching to native fields such as seconds or images.
1. Endpoints
| Operation | Method | Path |
|---|---|---|
| Submit a generation task | POST | /v1/video/generations |
| Submit a generation task (equivalent endpoint) | POST | /v1/videos |
| Query task status | GET | /v1/video/generations/{task_id} |
| Download the generated video | GET | /v1/videos/{task_id}/content |
Video generation is asynchronous:
- Submit a generation request.
- Retrieve the
task_idfrom the response. - Query the task until generation is complete.
- Download the generated video.
{BASE_URL} represents your platform’s API base URL. Authentication requirements, response schemas, task status values, and error codes are not included in the supplied specification. Refer to your platform’s documentation for these details.
2. Models and Capabilities
| Capability | MiniMax-H3 (Official) |
Self-Hosted Models |
|---|---|---|
| Model IDs | MiniMax-H3 |
minimax-h3-base, minimax-h3-base-fast, minimax-h3-mini |
| Duration | 4–15 seconds; default: 5 seconds | 5–15 seconds |
| Resolution | 768P (default), 2K |
720p only; 768P is converted to 720p |
| First and last frames | Supported | Supported |
| Reference images | Up to 9 images | Up to 9 images; maximum 20 MB per image |
| Reference videos | Up to 3 clips | Not supported |
| Reference audio | Up to 3 clips | Up to 3 clips; maximum 50 MB per clip |
| Output details | Includes native stereo audio | MP4, 24 fps |
The three self-hosted models use the same parameters and support the same capabilities. They differ in generation speed and pricing.
All reference media URLs must be publicly accessible from mainland China. Overseas direct links that cannot be accessed from mainland China are not supported.
3. Submit a Generation Task
POST {BASE_URL}/v1/video/generationsContent-Type: application/jsonEquivalent endpoint:
POST {BASE_URL}/v1/videos3.1 Unified Parameters
| Parameter | Type | Description |
|---|---|---|
model |
string | Required. The model ID. |
prompt |
string | Describes the video to generate. Required for self-hosted models, with a maximum length of 30,000 characters. |
duration |
number | Video duration in seconds. Official: 4–15. Self-hosted: 5–15. |
resolution |
string | Official: 768P or 2K. For self-hosted models, explicitly use 720p. |
aspect_ratio |
string | Output aspect ratio. See the rules below. |
image |
string | URL of a first-frame image or a single reference image. |
images |
array of strings | Image URLs; supported as a compatibility field. |
last_frame_image_url |
string | URL of the last-frame image. Use with image for first-and-last-frame generation. |
reference_image_urls |
array of strings | Reference image URLs. Up to 9 images. |
reference_video_urls |
array of strings | Reference video URLs. Official model only; up to 3 clips. |
reference_audio_urls |
array of strings | Reference audio URLs. Up to 3 clips. |
generation_mode |
string | Optional. t2v, i2v, or r2v. If omitted, the platform determines the mode from the supplied media. |
Explicitly set duration and resolution to avoid relying on model-specific defaults.
3.2 Aspect Ratio
Official model
- Text-to-video defaults to
16:9. adaptiveis recommended for first-and-last-frame or multi-reference generation.
Self-hosted models
Supported fixed aspect ratios:
16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 4:5, 5:4When aspect_ratio is set to adaptive, the platform omits the aspect-ratio field from the upstream request, allowing the upstream service to determine the output ratio.
3.3 Self-Hosted Options and Limits
| Parameter or Media | Rule |
|---|---|
prompt_optimization |
Optional boolean; defaults to false. When true, the upstream service rewrites the prompt using the official H3 prompt structure. |
| Reference images | Up to 9 images, no larger than 20 MB each. |
| Reference audio | Up to 3 clips, no larger than 50 MB each. |
| Multi-reference generation | Requires at least one reference image or one reference audio clip. |
| First and last frames | Use image and last_frame_image_url for the first and last frames, respectively. |
| Reference videos | Not supported. Do not send reference_video_urls. |
The platform does not reject out-of-range parameters for self-hosted models. Unsupported values such as a 4-second duration, 2K, or reference videos are forwarded upstream, where they may be ignored or cause an error. Billing is still provisionally deducted based on the duration declared in your request.
For self-hosted models, use 5–15 seconds, set resolution to 720p, and omit reference videos.
4. Request Examples
The following request bodies work with either submission endpoint. Replace the example media URLs with real URLs accessible from mainland China.
4.1 Official Model: Reference Image Generation
{ "model": "minimax-h3", "prompt": "The character from reference image 1 runs through a rainy street at night, with cinematic camera work.", "resolution": "768P", "duration": 8, "aspect_ratio": "16:9", "reference_image_urls": [ "https://example.com/character.png" ]}4.2 Self-Hosted: Text-to-Video
{ "model": "minimax-h3-base-fast", "prompt": "At dawn by the sea, a seagull glides over the water as the camera slowly moves forward in soft natural light.", "resolution": "720p", "duration": 8, "aspect_ratio": "16:9", "generation_mode": "t2v"}4.3 Self-Hosted: First-Frame Generation
{ "model": "minimax-h3-base", "prompt": "The character in the image runs through a rainy street at night, with cinematic camera work.", "resolution": "720p", "aspect_ratio": "16:9", "duration": 8, "image": "https://example.com/character.png"}4.4 Self-Hosted: First-and-Last-Frame Generation
{ "model": "minimax-h3-base", "prompt": "The character walks naturally from one end of the street to the other. Maintain a consistent appearance and use a smooth tracking shot.", "resolution": "720p", "duration": 8, "aspect_ratio": "adaptive", "image": "https://example.com/first-frame.png", "last_frame_image_url": "https://example.com/last-frame.png"}The platform passes the first and last frames to the upstream service in that order and automatically sets the upstream mode to fl2va.
4.5 Self-Hosted: Image and Audio References
{ "model": "minimax-h3-mini", "prompt": "Preserve the character's appearance from image 1 and make the character speak with the tone and rhythm of audio 1.", "resolution": "720p", "duration": 10, "reference_image_urls": [ "https://example.com/character.png" ], "reference_audio_urls": [ "https://example.com/voice.mp3" ]}4.6 Submit with cURL
Save a request body as request.json, then run:
curl --request POST "${BASE_URL}/v1/video/generations" \ --header "Content-Type: application/json" \ --data-binary @request.jsonAdd the authentication headers required by your platform.
5. Query Status and Download the Video
5.1 Query Task Status
Use the task ID returned by the submission request:
GET {BASE_URL}/v1/video/generations/{task_id}curl "${BASE_URL}/v1/video/generations/${TASK_ID}"Check the response to determine whether generation has completed. The exact status fields, status values, and error details depend on the platform’s response schema.
5.2 Download the Result
Once generation is complete:
GET {BASE_URL}/v1/videos/{task_id}/contentcurl --location --fail \ "${BASE_URL}/v1/videos/${TASK_ID}/content" \ --output video.mp4Include the platform’s required authentication information in both status and download requests.
6. Self-Hosted Parameter Translation
The platform performs these conversions automatically.
| Meaning | Unified Input | Upstream Field or Behavior |
|---|---|---|
| Duration | duration: 8 |
"seconds": "8" |
| Reference images or first frame | reference_image_urls, image, images, input_reference |
Merged in the listed order and deduplicated into images. |
| Last frame | last_frame_image_url |
Appended to images; automatically sets mode=fl2va. |
| Reference audio | reference_audio_urls |
audios |
| Generation mode | generation_mode |
Translated to mode; automatically inferred from supplied media if omitted. |
| Resolution | resolution: "768P" |
resolution: "720p"; no equivalent tier exists for 2K. |
| Adaptive aspect ratio | aspect_ratio: "adaptive" |
The aspect-ratio field is omitted. |
| Aspect-ratio alias | ratio |
Normalized to aspect_ratio. |
| Reference videos | reference_video_urls |
Unsupported; forwarded unchanged for upstream handling. |
For first-and-last-frame generation, use image and last_frame_image_url to specify the two frames. Avoid mixing additional image fields into the request, as merging them may affect image order.
7. Native Upstream Field Compatibility
Self-hosted models also accept native upstream fields directly. The platform leaves these native fields unchanged.
| Native Field | Accepted Format or Values |
|---|---|
seconds |
String or number, such as "8" or 8. |
images |
Array of image URLs. |
audios |
Array of audio URLs. |
mode |
t2va, i2va, fl2va, l2va, or ref2va. |
Example:
{ "model": "minimax-h3-base", "prompt": "The character in the image runs through a rainy street at night, with cinematic camera work.", "resolution": "720p", "aspect_ratio": "16:9", "seconds": "8", "images": [ "https://example.com/character.png" ], "mode": "i2va"}For new integrations, use the unified parameters. Conflict precedence between equivalent unified and native fields is not specified, so avoid sending both forms of the same setting—for example, duration and seconds—in one request.
