Skip to main content
Transcribe audio to text, fully compatible with OpenAI’s /v1/audio/transcriptions. Any OpenAI SDK can directly use it by pointing the base_url to https://api.acedata.cloud and replacing the key with your AceData Token. Supports standard complete responses and also supports gpt-transcribe’s SSE incremental transcription.
  • Request URL: POST https://api.acedata.cloud/v1/audio/transcriptions (alias POST /openai/audio/transcriptions)
  • Authentication: Request header Authorization: Bearer {token}
  • Request format: multipart/form-data
  • Billing: Charged by audio duration (see table below), rounded up to the nearest second.

Request Parameters

Which Model to Choose

Need subtitles or word-level timestamps → whisper-1; for other scenarios, recommend gpt-transcribe (more accurate and cheaper).

Example

Response:
Chinese audio is also supported without specifying the language:

Generate Subtitles

Set response_format to srt or vtt to directly obtain a usable subtitle file:
Returned content (Content-Type: text/plain):

Word-level Timestamps

To get the start and end times for each word, use verbose_json with timestamp_granularities[]=word:
Response:

Streaming Transcription

gpt-transcribe can return Content-Type: text/event-stream by using stream=true. The service will send OpenAI-compatible events as is: transcript.text.delta carries incremental text, transcript.text.done carries complete text and usage indicating normal completion.
Example of event stream:
Receiving transcript.text.done indicates normal completion. If processing fails after the stream is established, the connection will end after the event: error event; if the client disconnects actively, this processing will be canceled and will not continue generating in the background. whisper-1 will still return a normal non-streaming response even if stream=true is passed.

Using the Official SDK

Price

Charged based on the actual duration of the audio, with any duration less than 1 second counted as 1 second, and a maximum of 1 hour for a single request.

Notes

  • The maximum size for a single file is 25 MB. If it exceeds, please segment or compress it (usually lowering the bitrate is sufficient, as speech recognition does not have high audio quality requirements).
  • gpt-transcribe supports stream=true SSE; whisper-1 will ignore stream and return the complete result.
  • Parameters are consistent with OpenAI’s official /v1/audio/transcriptions, and the official SDK can be used by simply changing the base_url.
  • include[], chunking_strategy, known_speaker_names[], known_speaker_references[] belong to transcription models that we have not yet launched, passing them will return 400 instead of silently ignoring. Model-specific parameters (timestamp_granularities[] for whisper-1, languages[]/keywords[] for gpt-transcribe) will also return 400 when passed to unsupported models.
  • Requests can be time-consuming, it is recommended to set the client timeout to no less than 300 seconds.

Error Codes