Skip to main content

Dreamina Video Generation API

POST https://api.acedata.cloud/dreamina/videos Audio-driven talking-photo digital human generation (OmniHuman 1.5). Provide a portrait image and a driving audio to generate a video where the person speaks with synchronized lips.

Headers

Request parameters

Input guidance

  • Image: a clear, well-lit, front-facing portrait works best. The face should be unobstructed and reasonably large in frame.
  • Audio: mp3/wav, publicly reachable. Keep it under 60s (≤30s recommended for 1080p, ≤60s for 720p).
  • Both image_url and audio_url must be reachable from the public internet.

Response example

Async & retrieval

The endpoint runs synchronously by default and returns the finished video. For long jobs, use either async mode:
  • Pass callback_url — the endpoint returns a task_id immediately and POSTs the result to your URL when ready.
  • Pass async: true — the endpoint returns a task_id immediately; poll the result via POST /dreamina/tasks (free) by task_id or trace_id.
See the Dreamina Tasks API for the polling contract.

Error handling

Pricing

Billed by generated video duration — about ¥1/second at the largest package (e.g. a 10s video ≈ ¥10).