Skip to main content

Dreamina Digital Human Video Generation API

POST https://api.acedata.cloud/dreamina/videos Audio-driven digital human lip-sync video generation (OmniHuman 1.5). Provide a photo of a person and a piece of driving audio to generate a video of the person speaking with synchronized lip movements.

Request Headers

Request Parameters

Input Recommendations

  • Image: Clear, well-lit frontal portrait works best; face unobstructed, occupying a moderate proportion of the frame.
  • Audio: mp3/wav, must be publicly accessible. Recommended duration is within 60 seconds (1080p recommended ≤30 seconds, 720p ≤60 seconds).
  • Both image_url and audio_url must be publicly accessible.

Response Example

Asynchronous and Querying

The interface defaults to synchronous return of the final video. For longer tasks, you can use one of two asynchronous modes:
  • Provide callback_url: The interface immediately returns task_id, and the result is sent to this address after generation.
  • Provide async: true: The interface immediately returns task_id, then poll results via POST /dreamina/tasks (free) by task_id or trace_id.
For polling contracts, see Dreamina Tasks API.

Error Handling

Billing

Charged based on the duration of the generated video, maximum package approximately ¥1/second (e.g., a 10-second video costs about ¥10).