Skip to main content
Make an existing Kling video (5 seconds or 10 seconds) “speak” according to audio or text — that is, lip sync. Combined with the image2video feature of /kling/videos (to make a photo move), it can form a complete “talking photo / digital human presentation” workflow.
This API is a single-step convenient wrapper provided by AceDataCloud, designed for common audio/text-driven scenarios; it is not a field mirror of Kling’s official multi-step “Face Recognition → Advanced Lip Sync” API. Please refer to the parameter table on this page.
  • API Endpoint: POST https://api.acedata.cloud/kling/lip-sync
  • Request Format: application/json
  • Response Format: application/json
  • Billing: 2.45 Credits per successful call (fixed)

Request Headers

Request Parameters (Request Body)

Request Examples

1) Audio-driven (audio2video)

2) Text-driven (text2video)

Response Example (Synchronous Success)

Asynchronous Mode and Queries

When callback_url or async: true is passed, the API immediately returns task_id; afterwards you can:
  • Poll: POST /kling/tasks, body { "action": "retrieve", "id": "<task_id>" } (free)
  • Callback: After generation is complete, the result is POSTed to your callback_url

Complete Workflow: Talking Photo (image2video → lip-sync)

Error Response

Notes

  • video_id must be a Kling video generated within 30 days, and must be 5s or 10s; otherwise, use video_url to provide a video that meets the constraints.
  • The input video is recommended to feature a clear frontal face and a single person for the best lip-sync results.
  • The audio/text duration should match the video duration (audio must not exceed the video length).
  • Billing occurs upon success (2.45 Credits/request); parameter validation failures (4xx) are not billed.