Skip to main content
This article introduces the integration and use of the MiniMax H3 video generation API. This API supports text-to-video, first-and-last-frame control, and multimodal reference-based video generation, using a unified V2 multimodal content structure to create tasks.

Application Process

To use the MiniMax H3 video generation API, first go to the Ace Data Cloud Console to obtain your API Token and keep it for later use. If you have not logged in or registered yet, you will be automatically redirected to the login page and invited to register and log in. After completion, you will automatically return to the current page. One API Token can call all platform services; there is no need to apply separately for each service. Your first application includes free credits for a free trial; when credits are insufficient, you can top up your general balance in the Console.
📘 Full documentation: MiniMax H3 Video Generation API →
It is recommended to save the Token as an environment variable. Do not write it into source code or commit it to a version repository:

API Overview

  • Base URL: https://api.acedata.cloud
  • Endpoint: POST /minimax/videos
  • Authentication: Include authorization: Bearer {token} in the HTTP Header
  • Request Headers:
    • accept: application/json
    • content-type: application/json
  • Model (model): MiniMax-H3
  • Input Structure: Pass text, images, videos, and audio uniformly through content
  • Output Mode: By default, synchronously waits for generation to complete and returns the full task; when async: true or callback_url is passed, immediately returns task_id and trace_id
  • Result Query: Obtain status and completed videos through the MiniMax H3 Task Query API
  • Asynchronous Callback: Optional; receive the final task result through callback_url
You do not need to pass action to select a generation mode. The API automatically determines the purpose based on the material types and role in content.

Suitable Scenarios

Calling Process

When async is not passed by default, /minimax/videos waits for generation to complete and directly returns the full task. When you need to release the connection immediately, pass async: true or callback_url:
  1. Save the task_id and trace_id from the immediate response.
  2. When no callback is configured, call /minimax/tasks approximately every 10 seconds for querying.
  3. When task.status becomes succeeded, obtain the video from task.content.url.
  4. When the status is failed or cancelled, stop polling and read task.error.

Top-Level Request Parameters

The rules for ratio depend on the workflow:
  • Text-to-video: Required and cannot be adaptive.
  • First-frame, last-frame, or first-and-last-frame video: The aspect ratio is determined by the input image. It is recommended to omit it or pass adaptive.
  • Multimodal reference-based video generation: Can be omitted, with the default being adaptive; a fixed ratio can also be explicitly specified.
The API does not accept legacy or compatibility fields, such as prompt, image_urls, audio_urls, messages, and first_frame_image. When receiving errors for such parameters, delete the legacy fields and migrate to content; for example, change "prompt": "一只猫挥手" to "content": [{"type": "text", "text": "一只猫挥手"}]. Do not send both the new and legacy formats at the same time.

content Item Parameters

Each content item must have a type, and the remaining fields are determined by the type: Media addresses support three formats:
  • Publicly accessible HTTPS URLs, recommended for large files.
  • mm_file://{file_id}, referencing files that have already been uploaded or existing results.
  • Base64 data URIs for the corresponding media type. Base64 increases size by approximately one-third; please ensure the entire request body does not exceed 64 MB.

Material Specifications and Quantity Limits

Images, videos, and audio in multimodal reference scenarios total up to 12 files. The first/last frame scenario and the reference asset scenario are mutually exclusive: once reference_image, reference_video, or reference_audio is used, first_frame or last_frame can no longer be used, and vice versa.

Production-Grade Capability Showcase

The following are not concept images or placeholder assets, but the actual reference inputs and video outputs from MiniMax H3’s official production-grade capability samples. The three sets of examples respectively cover brand films, live-action narratives, and fashion e-commerce, suitable for evaluating the model’s most critical capabilities in commercial production. Here, “face capability” refers to character appearance consistency, facial details, and performance control in video generation; it is not identity recognition, face comparison, or face-swapping interfaces.

Premium Brand Film: Unifying Characters, Products, and Brand Assets

Production Goal: A 16:9 premium fashion brand film. Use a desert highway and a vintage car to establish a cool atmosphere, maintain the female protagonist’s appearance and the structure of the black handbag, and naturally incorporate the brand Logo into the ending. This example primarily tests cross-shot character consistency, product preservation, cinematic texture, and brand closure capabilities. Open or download the brand film directly The corresponding content structure:

Live-Action Vertical Short Drama: Face Consistency and Emotional Performance

Production Goal: A 15-second, 9:16 dark romance short drama trailer. Use the reference images of the male and female leads to lock in character appearance, and use the castle reference image to constrain the space; use medium close-ups and facial close-ups to portray eye contact confrontation, fear, restraint, and a sense of danger. This case is suitable for observing the stability of realistic facial features, micro-expressions, gaze relationships, and continuous performance. Open or download the realistic short drama directly The prompt should clearly specify the character relationship, emotions, and shot scale, rather than merely describing a “dialogue between a man and a woman”:

Fashion Eyewear Advertisement: Maintaining Facial Details and Product Structure

Production Goal: A 9:16 premium fashion eyewear advertisement. The full-body character image is responsible for body shape and runway walk, the face reference image is responsible for facial features and makeup, and the product image is responsible for the curved contours, lens reflections, temples, and cat-eye silhouette. This case simultaneously tests facial close-ups, multi-person consistency, wearing relationships, and product geometric structure. Open or download the fashion eyewear advertisement directly In product advertisements, the responsibilities of character references and product references should be clearly written separately in the prompt: character materials constrain the face, makeup, body shape, and temperament; product materials constrain the contours, materials, reflections, and wearing position. This is more stable than broadly writing “generate an eyewear advertisement.”

Text-to-Video

When there is only one text item, it is text-to-video. It is suitable for directly generating visuals from ideas, scripts, or shot descriptions. Prompts can be organized in the order of “subject + action + scene + camera + lighting + sound.”
The default synchronous mode returns the complete task after generation is finished:
If "async": true is added to the request, the API returns immediately:

First-Frame Image-to-Video

Mark the image as first_frame, and the model will begin generation from that image. It is suitable for naturally bringing posters, product images, character design images, and photography works to life.

End-Frame and First-and-Last-Frame Video

Providing only last_frame allows the model to naturally generate up to the specified frame; providing both first_frame and last_frame enables explicit control over the start and end points. Suitable for transitions, shape changes, growth processes, or before-and-after product comparisons.
The dimensions and aspect ratios of the first and last frames should be as consistent as possible, and differences in subject position, composition, and lighting should not be too large, making it easier to achieve a natural transition.

Multimodal Reference Video Generation

Reference materials can be used in combination: reference images control the appearance of characters or products, reference videos control actions and camera movement, and reference audio controls dialogue voice, music, or editing rhythm. The prompt should clearly state what each type of material should control, avoiding uploading materials without specifying their relationship.

Callback Notifications

Passing callback_url automatically enables asynchronous mode: the create API immediately returns task_id and trace_id, and POSTs the final result to that address after the task is completed, with the same structure as the task query response. The final status in the callback will be succeeded, failed, or cancelled. Even when using callbacks, it is recommended to save task_id for actively querying tasks or compensating for missed notifications.

Common Errors

task.status: succeeded in a synchronous response indicates that the video has been generated; asynchronous confirmation only indicates that the task has entered the queue. Charges apply only when a task ultimately succeeds; querying tasks is free and does not result in duplicate charges.

H3 Max

MiniMax-H3-Max supports 480P or 768P and integer durations of 5–15 seconds. Audio input is not charged additionally, the first 2 images are free, and each additional image is charged individually; reference videos are charged according to their actual input duration. This model does not support 2K.