> ## Documentation Index
> Fetch the complete documentation index at: https://docs.acedata.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI 음성 인식 API (/v1/audio/transcriptions)

> OpenAI generation API guide - Ace Data Cloud

오디오를 텍스트로 전사하며, **OpenAI의 `/v1/audio/transcriptions`와 완전히 호환됩니다**. 모든 OpenAI SDK는 `base_url`을 `https://api.acedata.cloud`로 설정하고, 키를 자신의 AceData Token으로 변경하면 바로 사용할 수 있습니다. 일반적인 전체 응답을 지원하며, `gpt-transcribe`의 SSE 증분 전사도 지원합니다.

* **요청 주소**: `POST https://api.acedata.cloud/v1/audio/transcriptions` (별칭 `POST /openai/audio/transcriptions`)
* **인증**: 요청 헤더 `Authorization: Bearer {token}`
* **요청 형식**: `multipart/form-data`
* **요금**: 오디오 길이에 따라 요금이 부과됩니다 (아래 표 참조), 1초 미만은 1초로 계산됩니다.

## 요청 매개변수

| 필드 | 유형 | 필수 | 설명 |
| - | - | - | - |
| `file` | file | 예 | 전사할 오디오 파일, 최대 25 MB. `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, `webm`을 지원합니다. |
| `model` | string | 아니오 | `whisper-1` (기본값) 또는 `gpt-transcribe`, 능력 차이는 아래 표를 참조하세요. |
| `language` | string | 아니오 | 오디오 언어, ISO-639-1 코드 (예: `zh`, `en`). 입력하면 정확도와 속도가 향상됩니다; 비워두면 자동으로 인식됩니다. |
| `prompt` | string | 아니오 | 작성 스타일을 유도하거나 고유 명사, 용어를 제공하여 인식 정확도를 높이는 데 사용되는 프롬프트입니다. |
| `response_format` | string | 아니오 | `whisper-1`: `json` (기본값), `text`, `srt`, `verbose_json`, `vtt`; `gpt-transcribe`: 오직 `json`, `text`. |
| `temperature` | number | 아니오 | 샘플링 온도 0–1, 기본값 0. |
| `timestamp_granularities[]` | array | 아니오 | 타임스탬프 세분화, `word` 또는 `segment`, `response_format=verbose_json`와 함께 사용해야 합니다. |
| `languages[]` | array | 아니오 | 후보 언어 (ISO-639-1), **오직 `gpt-transcribe`**. `language`와 상충하므로 동시에 전달하지 마세요. |
| `keywords[]` | array | 아니오 | 고유 명사/용어 프롬프트, \*\*오직 `gpt-transcribe`\*\*로, 브랜드명, 인명 인식 정확도를 크게 향상시킬 수 있습니다. |
| `stream` | boolean | 아니오 | `gpt-transcribe`가 `true`로 설정되면 SSE 증분 이벤트를 반환합니다; `whisper-1`은 이 매개변수를 무시하고 전체 결과를 반환합니다 (OpenAI 공식 동작과 일치). |

## 어떤 모델을 선택할까

| | `whisper-1` | `gpt-transcribe` |
| - | - | - |
| 가격 | \$0.0078 / 분 | **\$0.0059 / 분** (더 저렴) |
| 인식 정확도 | 양호 | **더 좋음**, 특히 브랜드명, 고유 명사에 대해 |
| 자막 출력 (`srt`/`vtt`) | ✅ | ❌ |
| 단어 수준 타임스탬프 | ✅ | ❌ |
| `languages[]` / `keywords[]` | ❌ | ✅ |
| 감지된 언어 반환 | `verbose_json` 필요 | 기본값으로 반환 |
| SSE 증분 반환 | ❌ (`stream` 무시됨) | ✅ |

**자막이나 단어 수준 타임스탬프가 필요하다면 → `whisper-1`; 나머지 경우는 `gpt-transcribe`를 추천합니다** (더 정확하고 더 저렴합니다).

## 예시

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1
```

반환:

```json theme={null}
{
  "text": "Ace Data Cloud Platform is testing the speech recognition endpoint. The quick brown fox jumps over the lazy dog."
}
```

중국어 오디오도 지원되며, 언어를 지정할 필요가 없습니다:

```json theme={null}
{
  "text": "欢迎使用 AceData Cloud 平台,我们正在测试语音识别接口,今天是 7 月 31 号。"
}
```

### 자막 생성

`response_format`을 `srt` 또는 `vtt`로 설정하면 바로 사용할 수 있는 자막 파일을 얻을 수 있습니다:

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1 \
  -F response_format=srt \
  -o subtitle.srt
```

반환 내용 (`Content-Type: text/plain`):

```
1
00:00:00,000 --> 00:00:03,800
Ace Data Cloud Platform is testing the speech recognition endpoint.

2
00:00:03,800 --> 00:00:06,280
The quick brown fox jumps over the lazy dog.
```

### 단어 수준 타임스탬프

각 단어의 시작과 끝 시간을 필요로 할 경우, `verbose_json`과 `timestamp_granularities[]=word`를 조합하여 사용합니다:

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F 'timestamp_granularities[]=word'
```

반환:

```json theme={null}
{
  "task": "transcribe",
  "language": "english",
  "duration": 6.29,
  "text": "Ace Data Cloud Platform is testing the speech recognition endpoint. The quick brown fox jumps over the lazy dog.",
  "words": [
    {
      "word": "Ace",
      "start": 0.0,
      "end": 0.32
    },
    {
      "word": "Data",
      "start": 0.32,
      "end": 0.54
    },
    {
      "word": "Cloud",
      "start": 0.54,
      "end": 0.86
    }
  ]
}
```

### 스트리밍 전사

`gpt-transcribe`는 `stream=true`를 통해 `Content-Type: text/event-stream`을 반환할 수 있습니다. 서비스는 OpenAI 호환 이벤트를 그대로 전송합니다:
`transcript.text.delta`는 증분 텍스트를 포함하고, `transcript.text.done`은 전체 텍스트와 `usage`를 포함하여 정상 완료를 나타냅니다.

```shell theme={null}
curl -N -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=gpt-transcribe \
  -F stream=true
```

이벤트 스트림 예시:

```text theme={null}
data: {"type":"transcript.text.delta","delta":"Hello"}

data: {"type":"transcript.text.done","text":"Hello world","usage":{"type":"tokens","input_tokens":14,"output_tokens":3,"total_tokens":17}}
```

`transcript.text.done`을 수신해야 정상 완료를 나타냅니다. 스트림이 설정된 후 처리에 실패하면 연결은 `event: error` 이벤트 후 종료됩니다; 클라이언트가 적극적으로 연결을 끊으면 이번 처리가 취소되며 백그라운드에서 계속 생성되지 않습니다. `whisper-1`은 `stream=true`를 전달하더라도 여전히 일반 비스트리밍 응답으로 반환됩니다.

### 공식 SDK 사용

```python theme={null}
from openai import OpenAI

client = OpenAI(base_url="https://api.acedata.cloud/v1", api_key="{token}")
with open("audio.mp3", "rb") as f:
    result = client.audio.transcriptions.create(model="whisper-1", file=f)
print(result.text)
```

## 가격

| 모델 | 본 플랫폼가 |
| - | - |
| `whisper-1` | \$0.0078 / 분 |
| `gpt-transcribe` | \$0.0059 / 분 |

> 음성의 실제 길이에 따라 요금이 부과되며, 1초 미만은 1초로 계산하고, 단일 최대 길이는 1시간으로 제한됩니다.

## 주의사항

* 단일 파일 최대 **25 MB**. 초과 시 먼저 분할하거나 압축하십시오(비트 전송률을 낮추는 것으로 충분하며, 음성 인식은 음질 요구가 높지 않습니다).
* `gpt-transcribe`는 `stream=true` SSE를 지원합니다; `whisper-1`은 `stream`을 무시하고 전체 결과를 반환합니다.
* 매개변수는 OpenAI 공식 `/v1/audio/transcriptions`와 일치하며, 공식 SDK는 `base_url`만 변경하면 사용할 수 있습니다.
* `include[]`、`chunking_strategy`、`known_speaker_names[]`、`known_speaker_references[]`는 우리가 출시하지 않은
  전사 모델에 해당하며, 전달 시 400을 반환하고 조용히 무시되지 않습니다. 모델 전용 매개변수(`timestamp_granularities[]`는 `whisper-1`에,
  `languages[]`/`keywords[]`는 `gpt-transcribe`에 해당)도 지원하지 않는 모델에 전달 시 400을 반환합니다.
* 요청이 다소 시간이 소요되므로 클라이언트의 타임아웃 설정은 300초 이상으로 설정하는 것이 좋습니다.

## 오류 코드

| 상태 코드 | 코드 | 설명 |
| - | - | - |
| 400 | `bad_request` | `file`이 제공되지 않았거나, 파일을 해석할 수 없거나, 매개변수가 잘못되었습니다(`model`/`response_format`의 값이 지원되지 않음, `temperature`가 0–1을 초과, `timestamp_granularities[]`가 `verbose_json`과 함께 사용되지 않음, `whisper-1`이 지원하지 않는 매개변수가 전달됨). |
| 401 | `authentication_failed` | 토큰이 유효하지 않습니다. |
| 403 | `used_up` | 잔액이 부족합니다. |
| 413 | `request_too_large` | 오디오 파일이 25 MB 한도를 초과했습니다. |
| 429 | `too_many_requests` | 요청이 너무 빈번합니다. 잠시 후 다시 시도하십시오. |
| 500 | `api_error` | 서비스 내부 오류입니다. 잠시 후 다시 시도하십시오. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.