> ## Documentation Index
> Fetch the complete documentation index at: https://docs.acedata.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI 語音識別 API（/v1/audio/transcriptions）

> OpenAI generation 整合指南 - Ace Data Cloud

將音頻轉寫為文字，**完全兼容 OpenAI 的 `/v1/audio/transcriptions`**。任何 OpenAI SDK 只需把 `base_url` 指向 `https://api.acedata.cloud`、把密鑰換成你的 AceData Token 即可直接使用。支持普通完整響應，也支持 `gpt-transcribe` 的 SSE 增量轉寫。

* **請求地址**：`POST https://api.acedata.cloud/v1/audio/transcriptions`（別名 `POST /openai/audio/transcriptions`）
* **鑑權**：請求頭 `Authorization: Bearer {token}`
* **請求格式**：`multipart/form-data`
* **計費**：按音頻時長計費（見下表），不足 1 秒按 1 秒計。

## 請求參數

| 字段 | 類型 | 必填 | 說明 |
| - | - | - | - |
| `file` | file | 是 | 待轉寫的音頻文件，最大 25 MB。支持 `flac`、`mp3`、`mp4`、`mpeg`、`mpga`、`m4a`、`ogg`、`wav`、`webm`。 |
| `model` | string | 否 | `whisper-1`（默認）或 `gpt-transcribe`，能力差異見下表。 |
| `language` | string | 否 | 音頻語種，ISO-639-1 代碼（如 `zh`、`en`）。填寫可提升準確率與速度；留空則自動識別。 |
| `prompt` | string | 否 | 提示詞，用於引導書寫風格，或提供專有名詞、術語以提升識別準確率。 |
| `response_format` | string | 否 | `whisper-1`：`json`（默認）、`text`、`srt`、`verbose_json`、`vtt`；`gpt-transcribe`：僅 `json`、`text`。 |
| `temperature` | number | 否 | 采樣溫度 0–1，默認 0。 |
| `timestamp_granularities[]` | array | 否 | 時間戳粒度，`word` 或 `segment`，需配合 `response_format=verbose_json` 使用。 |
| `languages[]` | array | 否 | 候選語種（ISO-639-1），**僅 `gpt-transcribe`**。與 `language` 互斥，不要同時傳。 |
| `keywords[]` | array | 否 | 專有名詞/術語提示，**僅 `gpt-transcribe`**，可顯著提升品牌名、人名的識別準確率。 |
| `stream` | boolean | 否 | `gpt-transcribe` 設為 `true` 時返回 SSE 增量事件；`whisper-1` 會忽略該參數並返回完整結果（與 OpenAI 官方行為一致）。 |

## 選哪個模型

| | `whisper-1` | `gpt-transcribe` |
| - | - | - |
| 價格 | \$0.0078 / 分鐘 | **\$0.0059 / 分鐘**（更便宜） |
| 識別準確率 | 良好 | **更好**，尤其是品牌名、專有名詞 |
| 字幕輸出（`srt`/`vtt`） | ✅ | ❌ |
| 詞級時間戳 | ✅ | ❌ |
| `languages[]` / `keywords[]` | ❌ | ✅ |
| 返回檢測到的語種 | 需 `verbose_json` | 默認返回 |
| SSE 增量返回 | ❌（`stream` 被忽略） | ✅ |

**需要字幕或詞級時間戳 → `whisper-1`；其餘場景推薦 `gpt-transcribe`**（更準且更便宜）。

## 示例

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1
```

返回：

```json theme={null}
{
  "text": "Ace Data Cloud Platform is testing the speech recognition endpoint. The quick brown fox jumps over the lazy dog."
}
```

中文音頻同樣支持，無需指定語種：

```json theme={null}
{
  "text": "歡迎使用 AceData Cloud 平台,我們正在測試語音識別接口,今天是 7 月 31 號。"
}
```

### 生成字幕

把 `response_format` 設為 `srt` 或 `vtt`，直接得到可用的字幕文件：

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1 \
  -F response_format=srt \
  -o subtitle.srt
```

返回內容（`Content-Type: text/plain`）：

```
1
00:00:00,000 --> 00:00:03,800
Ace Data Cloud Platform is testing the speech recognition endpoint.

2
00:00:03,800 --> 00:00:06,280
The quick brown fox jumps over the lazy dog.
```

### 詞級時間戳

需要每個詞的起止時間時，用 `verbose_json` 搭配 `timestamp_granularities[]=word`：

```shell theme={null}
curl -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=whisper-1 \
  -F response_format=verbose_json \
  -F 'timestamp_granularities[]=word'
```

返回：

```json theme={null}
{
  "task": "transcribe",
  "language": "english",
  "duration": 6.29,
  "text": "Ace Data Cloud Platform is testing the speech recognition endpoint. The quick brown fox jumps over the lazy dog.",
  "words": [
    {
      "word": "Ace",
      "start": 0.0,
      "end": 0.32
    },
    {
      "word": "Data",
      "start": 0.32,
      "end": 0.54
    },
    {
      "word": "Cloud",
      "start": 0.54,
      "end": 0.86
    }
  ]
}
```

### 流式轉寫

`gpt-transcribe` 可通過 `stream=true` 返回 `Content-Type: text/event-stream`。服務會原樣發送 OpenAI 兼容事件：
`transcript.text.delta` 攜帶增量文字，`transcript.text.done` 攜帶完整文字和 `usage` 並表示正常完成。

```shell theme={null}
curl -N -X POST 'https://api.acedata.cloud/v1/audio/transcriptions' \
  -H 'authorization: Bearer {token}' \
  -F file=@audio.mp3 \
  -F model=gpt-transcribe \
  -F stream=true
```

事件流示例：

```text theme={null}
data: {"type":"transcript.text.delta","delta":"Hello"}

data: {"type":"transcript.text.done","text":"Hello world","usage":{"type":"tokens","input_tokens":14,"output_tokens":3,"total_tokens":17}}
```

收到 `transcript.text.done` 才表示正常完成。若流建立後處理失敗，連接會在 `event: error` 事件後結束；客戶端主動斷開會取消本次處理，不會繼續在後台生成。`whisper-1` 即使傳入 `stream=true` 也仍按普通非流式響應返回。

### 使用官方 SDK

```python theme={null}
from openai import OpenAI

client = OpenAI(base_url="https://api.acedata.cloud/v1", api_key="{token}")
with open("audio.mp3", "rb") as f:
    result = client.audio.transcriptions.create(model="whisper-1", file=f)
print(result.text)
```

## 價格

| 模型 | 本平台價 |
| - | - |
| `whisper-1` | \$0.0078 / 分鐘 |
| `gpt-transcribe` | \$0.0059 / 分鐘 |

> 按音頻實際時長計費，不足 1 秒按 1 秒計，單次最長按 1 小時封頂。

## 注意事項

* 單個文件最大 **25 MB**。超過請先分段或壓縮（降低碼率通常即可，語音識別對音質要求不高）。
* `gpt-transcribe` 支持 `stream=true` SSE；`whisper-1` 會忽略 `stream` 並返回完整結果。
* 參數與 OpenAI 官方 `/v1/audio/transcriptions` 保持一致，官方 SDK 只需改 `base_url` 即可使用。
* `include[]`、`chunking_strategy`、`known_speaker_names[]`、`known_speaker_references[]` 屬於我們未上架的
  轉寫模型，傳入會返回 400 而不是靜默忽略。模型專屬參數（`timestamp_granularities[]` 之於 `whisper-1`、
  `languages[]`/`keywords[]` 之於 `gpt-transcribe`）傳給不支持的模型時同樣返回 400。
* 請求較為耗時，建議客戶端超時設置不低於 300 秒。

## 錯誤碼

| 狀態碼 | code | 說明 |
| - | - | - |
| 400 | `bad_request` | 未提供 `file`、文件無法解析，或參數非法（`model`/`response_format` 取值不支持、`temperature` 超出 0–1、`timestamp_granularities[]` 未配合 `verbose_json`、傳入了 `whisper-1` 不支持的參數）。 |
| 401 | `authentication_failed` | token 無效。 |
| 403 | `used_up` | 餘額不足。 |
| 413 | `request_too_large` | 音頻文件超過 25 MB 上限。 |
| 429 | `too_many_requests` | 請求過於頻繁，請稍後重試。 |
| 500 | `api_error` | 服務內部錯誤，請稍後重試。 |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.