Veo Video Generation API

Official Gemini Veo asynchronous protocol: submit a video generation task → poll the operation → retrieve the video result

POST/v1beta/models/{model}:predictLongRunning

Veo models use the official Gemini Veo protocol and are transparently proxied by the platform. The request body uses instances and parameters. The submission returns an operation; poll it until done=true to retrieve the result. Refer to List Models for currently available model names, such as veo-3.1-generate-preview and veo-3.1-fast-generate-preview.

Authorizations

AuthorizationstringRequired

Bearer Token authentication. No Google API Key or GCP credentials are required. Get an API Key from the API Key management page.

Authorization: Bearer YOUR_API_KEY

Submit a Generation Task

POST /v1beta/models/{model}:predictLongRunning

Path parameters:

  • model (required): Veo model name. Use a model returned by the List Models endpoint.

The request body follows the official Gemini Veo format:

  • instances (object[], required): Generation input
    • prompt (string, required): Description of the video content, motion, camera work, style, and sound
    • image (object, optional): First-frame image for image-to-video generation, where supported by the model
      • bytesBase64Encoded (string, required): Base64-encoded image content without a Data URL prefix
      • mimeType (string, required): Image MIME type, such as image/png or image/jpeg
  • parameters (object, optional): Generation parameters
    • sampleCount (integer, optional): Number of videos to generate
    • durationSeconds (integer, optional): Video duration in seconds
    • aspectRatio (string, optional): Frame aspect ratio, such as 16:9 or 9:16
    • resolution (string, optional): Output resolution, such as 720p or 1080p
    • negativePrompt (string, optional): Content that should not appear in the video
    • personGeneration (string, optional): Person-generation policy
    • seed (integer, optional): Random seed
    • enhancePrompt (boolean, optional): Whether to enhance the prompt
    • generateAudio (boolean, optional): Whether to generate audio; only available for models with audio support

Supported parameters and value ranges vary between Veo models. The platform passes official protocol fields through unchanged. The model returns an error when a field or value is unsupported.

Text-to-Video Request Example

curl --request POST \  --url https://api.maitoken.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning \  --header 'Authorization: Bearer <token>' \  --header 'Content-Type: application/json' \  --data '{    "instances": [      {        "prompt": "A cinematic aerial shot of ocean waves crashing against rocks at sunset, with the camera slowly moving forward"      }    ],    "parameters": {      "sampleCount": 1,      "durationSeconds": 8,      "aspectRatio": "16:9",      "resolution": "720p",      "generateAudio": true    }  }'

Image-to-Video Request Example

curl --request POST \  --url https://api.maitoken.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning \  --header 'Authorization: Bearer <token>' \  --header 'Content-Type: application/json' \  --data '{    "instances": [      {        "prompt": "Keep the main subject consistent, slowly move the camera forward, and let the clouds in the background flow naturally",        "image": {          "bytesBase64Encoded": "IMAGE_BASE64",          "mimeType": "image/png"        }      }    ],    "parameters": {      "sampleCount": 1,      "durationSeconds": 8,      "aspectRatio": "16:9",      "resolution": "720p"    }  }'

Query Task Status

GET /v1beta/{operation_name}

Use the complete name returned by the submission as operation_name. Do not replace the model name or use only the task ID.

Processing Response

The task is still processing while done is false or until the response includes done=true:

{  "name": "models/veo-3.1-generate-preview/operations/OPERATION_ID",  "done": false}

Successful Response

When the task completes, done=true and the generated results are available in response.generateVideoResponse.generatedSamples[]:

{  "name": "models/veo-3.1-generate-preview/operations/OPERATION_ID",  "done": true,  "response": {    "generateVideoResponse": {      "generatedSamples": [        {          "video": {            "uri": "https://api.maitoken.com/v1beta/files/FILE_ID:download?alt=media",            "encoding": "video/mp4"          }        }      ]    }  }}

Depending on the channel configuration, video contains one of the following result formats:

  • uri: Video download URL
  • encodedVideo: Base64-encoded video content
  • encoding: Video MIME type, such as video/mp4

Failed Response

When the task fails, done=true and error details are returned in the top-level error field:

{  "name": "models/veo-3.1-generate-preview/operations/OPERATION_ID",  "done": true,  "error": {    "code": 3,    "message": "Invalid request parameters"  }}

Download the Video

When the result contains video.uri, download it using the complete returned URL and the same API Key used to submit the task:

curl --location \  --url 'VIDEO_URI' \  --header 'Authorization: Bearer <token>' \  --output veo.mp4

When the result contains video.encodedVideo, decode it as standard Base64 content and save it as a video file.

Notes

  • Task queries and file downloads must use the same API Key that was used to submit the task.
  • Poll every 5–10 seconds until the response returns done=true.
  • Use exponential backoff for 429 and 5xx responses to avoid excessive polling.
  • A timed-out submission request does not necessarily mean task creation failed. Retrying before confirming the outcome may create duplicate tasks and incur duplicate charges.
  • After obtaining an operation name, always query the original task instead of submitting another one while waiting for the result.
  • A client-side timeout does not mean the server-side task failed. You can continue querying it later with the original operation.
  • Video URLs may expire. Download and store the result promptly after the task succeeds.
  • Explicitly provide durationSeconds, resolution, and generateAudio instead of relying on implicit defaults that may differ between model versions.
  • Supported durations, aspect ratios, resolutions, audio capabilities, and image-to-video features depend on the model currently available to your account.