> ## Documentation Index
> Fetch the complete documentation index at: https://docs.acedata.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# GLM 채팅 완료 API 신청 및 사용

> GLM API guide - Ace Data Cloud

GLM(General Language Model)은 지푸 AI(Zhipu AI / Z.ai)가 출시한 차세대 대형 언어 모델 시리즈로, 강력한 중문 및 영문 이해 및 생성 능력을 갖추고 있으며, 중국어 환경, 코드 생성, 추론 및 다중 회차 대화 등 작업에서 뛰어난 성능을 보입니다. GLM-5.3, GLM-5.2, GLM-4.7 등 차세대 모델은 긴 문맥, 도구 호출 및 코드 작업에서 많은 최적화를 이루어냈으며, 스마트 Q\&A, 콘텐츠 창작, 코드 보조, 고객 서비스 로봇 등 다양한 분야에 널리 활용될 수 있습니다.

이 문서는 GLM 채팅 완료 API의 사용 프로세스를 주로 소개하며, 이를 통해 통합된 OpenAI 호환 인터페이스를 통해 GLM 시리즈 모델을 쉽게 호출할 수 있습니다.

## 신청 프로세스

GLM 채팅 완료 API를 사용하려면 먼저 [Ace Data Cloud 콘솔](https://platform.acedata.cloud/console/applications)에서 API 토큰을 받아야 하며, 이를 백업용으로 보관하십시오.

![](https://cdn.acedata.cloud/dvc3cg.jpg)

로그인 또는 등록이 되어 있지 않은 경우, 자동으로 로그인 페이지로 리디렉션되어 등록 및 로그인을 요청합니다. 완료 후 현재 페이지로 자동으로 돌아옵니다.

**하나의 API 토큰으로 플랫폼의 모든 서비스를 호출할 수 있으며, 각 서비스마다 별도로 신청할 필요가 없습니다.** 처음 신청 시 무료 할당량이 제공되어 무료로 체험할 수 있으며, 할당량이 부족할 경우 [콘솔](https://platform.acedata.cloud/console/coin)에서 일반 잔액을 충전할 수 있습니다.

> 📘 전체 문서: [GLM 채팅 완료 API →](https://platform.acedata.cloud/documents/glm-chat-completions)

## 기본 사용

GLM 채팅 완료 API의 요청 주소는 `https://api.acedata.cloud/glm/chat/completions`이며, Bearer Token으로 인증합니다. 요청 본문은 OpenAI 채팅 완료 프로토콜과 호환됩니다.

이 인터페이스를 처음 사용할 때, 최소한 세 가지 내용을 입력해야 합니다:

* `authorization`: 드롭다운 목록에서 Bearer Token을 선택하면 됩니다.
* `model`: 호출할 GLM 모델을 선택합니다. 현재 지원되는 모델은 다음과 같습니다:
  * `glm-5.3`: 최신 플래그십 모델로, 1M 문맥 및 최대 128K 출력을 지원하며, 복잡한 추론, 코드 및 에이전트 작업에 적합합니다. 추론은 항상 활성화되어 있으며, `reasoning_effort`를 통해 `low`, `high` 또는 `max`를 선택할 수 있습니다.
  * `glm-5.2`: 이전 세대 플래그십 모델로, 종합 능력이 뛰어납니다.
  * `glm-5.1`: 성숙한 플래그십 모델로, 일반 복잡한 작업에 적합합니다.
  * `glm-4.7`: 추론, 도구 호출 및 코드 작업에서 우수한 성능을 보입니다.
  * `glm-4.6`: 일반 대화 모델로, 효과와 비용의 균형을 맞춥니다.
  * `glm-3-turbo`: 고전 대화 모델로, 일반 텍스트 생성 작업에 적합합니다.
* `messages`: 프롬프트 배열로, 각 메시지는 `role`과 `content`를 포함하며, `role`은 `user`, `assistant`, `system` 세 가지 역할을 지원합니다.

자주 사용하는 선택적 매개변수:

* `max_tokens`: 단일 응답의 최대 토큰 수를 제한합니다.
* `temperature`: 생성의 무작위성으로, 0-2 사이의 값이며, 값이 클수록 분산됩니다.
* `top_p`: 핵 샘플링 매개변수로, 후보 토큰의 누적 확률 임계값을 제어합니다.
* `n`: 한 번에 생성할 후보 응답 수입니다.
* `stream`: 스트리밍 응답을 활성화할지 여부로, 기본값은 `false`입니다.
* `stop`: 사용자 정의 중지 시퀀스입니다.

다음은 가장 간단한 Python 호출 예제입니다:

```python theme={null}
import requests

url = "https://api.acedata.cloud/glm/chat/completions"

headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}

payload = {
    "model": "glm-5.2",
    "messages": [
        {"role": "user", "content": "hello"}
    ]
}

response = requests.post(url, json=payload, headers=headers)
print(response.text)
```

호출 후, 반환된 결과는 다음과 같습니다:

```json theme={null}
{
  "id": "msg_202604262252030313862701a04e33",
  "model": "glm-5.2",
  "object": "chat.completion",
  "created": 1777215124,
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! 👋 How can I assist you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 10,
    "completion_tokens": 23,
    "total_tokens": 33
  }
}
```

반환된 결과의 주요 필드 설명은 다음과 같습니다:

* `id`: 이번 대화 작업의 고유 ID입니다.
* `created`: 이번 대화 작업의 생성 시간(Unix 타임스탬프, 초)입니다.
* `model`: 실제 호출된 GLM 모델 이름입니다.
* `choices`: 모델이 생성한 응답 목록입니다. `choices[i].message.content`는 모델의 구체적인 응답 텍스트이며, `finish_reason`은 종료 이유를 나타냅니다(`stop`, `length`, `tool_calls`, `content_filter` 등).
* `usage`: 이번 요청의 토큰 사용량 통계로, `prompt_tokens`, `completion_tokens`, `total_tokens`를 포함합니다.

## 스트리밍 응답

이 인터페이스는 스트리밍 응답(서버 전송 이벤트)을 지원하며, 이는 웹 페이지 통합에 매우 유용하여 웹 페이지에서 글자 단위로 표시하는 효과를 구현할 수 있습니다.

스트리밍 응답을 원할 경우, 요청 본문에서 `stream` 매개변수를 `true`로 설정하면 됩니다.

Python 샘플 호출 코드:

```python theme={null}
import requests

url = "https://api.acedata.cloud/glm/chat/completions"

headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}

payload = {
    "model": "glm-4.7",
    "messages": [{"role": "user", "content": "hi"}],
    "stream": True
}

response = requests.post(url, json=payload, headers=headers, stream=True)
for line in response.iter_lines():
    if line:
        print(line.decode("utf-8"))
```

출력 효과는 다음과 같습니다(발췌):

```text theme={null}
data: {"id": "msg_2026042622521271f765bbc3734ce1", "object": "chat.completion.chunk", "created": 1777215133, "model": "glm-4.7", "choices": [{"delta": {"content": "", "role": "assistant"}, "finish_reason": null, "index": 0}], "usage": null}

data: {"id": "msg_2026042622521271f765bbc3734ce1", "object": "chat.completion.chunk", "created": 1777215133, "model": "glm-4.7", "choices": [{"delta": {"content": "안녕하세요! 제가 도와드릴 수 있는 것이 있나요?"}, "finish_reason": null, "index": 0}], "usage": null}

data: {"id": "msg_2026042622521271f765bbc3734ce1", "object": "chat.completion.chunk", "created": 1777215133, "model": "glm-4.7", "choices": [{"delta": {"content": "도와드릴까요?"}, "finish_reason": null, "index": 0}], "usage": null}

data: {"id": "msg_2026042622521271f765bbc3734ce1", "object": "chat.completion.chunk", "created": 1777215133, "model": "glm-4.7", "choices": [{"delta": {}, "finish_reason": "stop", "index": 0}], "usage": null}

data: {"id": "msg_2026042622521271f765bbc3734ce1", "object": "chat.completion.chunk", "created": 1777215133, "model": "glm-4.7", "choices": [], "usage": {"prompt_tokens": 1420, "completion_tokens": 18, "total_tokens": 1438}}

data: [DONE]
```

위의 내용을 보면, 응답 안에 여러 개의 `data`가 있으며, 각 `data`는 하나의 증분 조각을 포함하고 있습니다. `choices[i].delta.content`는 현재 청크에 추가된 텍스트 조각이며, 이 조각들을 연결하여 전체 응답을 형성할 수 있습니다. `data` 내용이 `[DONE]`일 때는 스트리밍 응답이 끝났음을 나타냅니다. 마지막으로 `usage`가 있는 청크는 이번 요청의 토큰 사용량을 요약합니다.

JavaScript(노드.js) 샘플:

```javascript theme={null}
const options = {
  method: "POST",
  headers: {
    accept: "application/json",
    authorization: "Bearer {token}",
    "content-type": "application/json"
  },
  body: JSON.stringify({
    model: "glm-4.7",
    messages: [{ role: "user", content: "안녕하세요" }],
    stream: true
  })
};

const response = await fetch("https://api.acedata.cloud/glm/chat/completions", options);
const reader = response.body.getReader();
const decoder = new TextDecoder("utf-8");
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  process.stdout.write(decoder.decode(value));
}
```

Java 샘플 코드:

```java theme={null}
JSONObject jsonObject = new JSONObject();
jsonObject.put("model", "glm-4.7");
jsonObject.put("messages", new JSONArray().put(new JSONObject().put("role", "user").put("content", "안녕하세요")));
jsonObject.put("stream", true);
MediaType mediaType = MediaType.parse("application/json; charset=utf-8");
RequestBody body = RequestBody.create(jsonObject.toString(), mediaType);
Request request = new Request.Builder()
  .url("https://api.acedata.cloud/glm/chat/completions")
  .post(body)
  .addHeader("accept", "application/json")
  .addHeader("authorization", "Bearer {token}")
  .addHeader("content-type", "application/json")
  .build();

OkHttpClient client = new OkHttpClient();
Response response = client.newCall(request).execute();
System.out.println(response.body().string());
```

다른 언어는 별도로 수정할 수 있으며, 원리는 모두 동일합니다.

## 다중 대화

다중 대화 기능을 구현하려면, 이전 대화를 순서대로 `messages` 배열에 넣고 `user`와 `assistant`가 번갈아 나타나는 순서를 유지해야 합니다.

Python 샘플 호출 코드:

```python theme={null}
import requests

url = "https://api.acedata.cloud/glm/chat/completions"

headers = {
    "accept": "application/json",
    "authorization": "Bearer {token}",
    "content-type": "application/json"
}

payload = {
    "model": "glm-4.7",
    "messages": [
        {"role": "user", "content": "안녕하세요"},
        {"role": "assistant", "content": "안녕하세요! 오늘 무엇을 도와드릴까요?"},
        {"role": "user", "content": "방금 제가 뭐라고 했죠?"}
    ]
}

response = requests.post(url, json=payload, headers=headers)
print(response.text)
```

여러 개의 질문어를 업로드하면 쉽게 다중 대화를 구현할 수 있으며, 다음과 같은 응답을 얻을 수 있습니다:

```json theme={null}
{
  "id": "msg_20260426225208b95324e9945a48d3",
  "model": "glm-4.7",
  "object": "chat.completion",
  "created": 1777215128,
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "당신이 말한 것은: **\"안녕하세요\"** 😊\n\n다른 것이 필요하시면 알려주세요!"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 48,
    "completion_tokens": 37,
    "total_tokens": 85
  }
}
```

`choices`에 포함된 정보는 기본 사용과 일치하며, 모델은 전체 대화 기록을 기반으로 응답을 제공하여 다중 맥락 상호작용을 지원합니다.

## 시스템 프롬프트(System Prompt)

`messages`의 시작 부분에 `role`이 `system`인 메시지를 추가하여 모델의 역할, 스타일 또는 행동을 제약할 수 있습니다:

```python theme={null}
payload = {
    "model": "glm-4.7",
    "messages": [
        {"role": "system", "content": "당신은 숙련된 한국어 작문 도우미입니다. 간결하고 전문적인 어조로 응답하세요."},
        {"role": "user", "content": "GLM 모델에 대해 세 문장으로 소개해 주세요."}
    ]
}
```

## 도구 호출(Function Calling)

GLM 모델은 OpenAI 호환의 Function Calling을 지원하며, `tools` 매개변수를 통해 호출 가능한 함수를 선언할 수 있습니다. 모델은 필요할 때 `choices[i].message.tool_calls`에서 구조화된 함수 호출 정보를 반환합니다.

```python theme={null}
payload = {
    "model": "glm-4.7",
    "messages": [
        {"role": "user", "content": "오늘 베이징 날씨 어때요?"}
    ],
    "tools": [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "지정된 도시의 날씨를 조회합니다.",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "city": {"type": "string", "description": "도시 이름"}
                    },
                    "required": ["city"]
                }
            }
        }
    ]
}
```

모델이 도구를 호출하기로 결정하면, 반환 결과의 `finish_reason`이 `tool_calls`로 변경되며, `message.tool_calls`에서 함수 이름과 JSON 문자열 형식의 매개변수를 제공합니다. 해당 함수를 실행하고 결과를 `role`이 `tool`인 메시지로 모델에 다시 전달하여 완전한 도구 호출 루프를 완료할 수 있습니다.

## 모델 선택 제안

| 모델 | 적합한 장면 |
| - | - |
| `glm-5.3` | 최신 플래그십, 1M 컨텍스트, 최장 128K 출력, 복잡한 추론, 코드 및 에이전트 작업에 추천 |
| `glm-5.2` | 이전 세대 플래그십, 복잡한 추론, 코드 및 에이전트 작업에 적합 |
| `glm-5.1` | 성숙한 플래그십, 복잡한 추론, 긴 문서 분석에 적합 |
| `glm-4.7` | 도구 호출, 코드 생성, 에이전트 오케스트레이션 등의 작업 |
| `glm-4.6` | 일반 대화, 콘텐츠 창작의 균형 잡힌 선택 |
| `glm-3-turbo` | 일반 텍스트 생성 작업, 비용에 민감한 장면 |

## 오류 처리

API를 호출할 때 오류가 발생하면, API는 해당 오류 코드와 정보를 반환합니다. 예를 들어:

* `400 token_mismatched`：요청 매개변수가 누락되었거나 유효하지 않음.
* `400 api_not_implemented`：지원되지 않는 매개변수나 모델을 사용함.
* `401 invalid_token`：권한 없음, Bearer Token이 누락되었거나 만료됨.
* `429 too_many_requests`：빈도 제한이 발생함, 나중에 다시 시도하십시오.
* `500 api_error`：서버 내부 오류 또는 상류가 일시적으로 사용할 수 없음.

### 오류 응답 예시

```json theme={null}
{
  "trace_id": "69ea9bcf-c5da-41a3-be97-c80912a08523",
  "error": {
    "code": "api_error",
    "message": "서비스가 일시적으로 사용할 수 없습니다. 나중에 다시 시도하십시오."
  }
}
```

`api_error`가 반환되고 메시지가 `서비스가 일시적으로 사용할 수 없습니다. 나중에 다시 시도하십시오.`일 때, 일반적으로 상류 GLM 서비스가 일시적으로 사용할 수 없음을 나타내며, 지수 백오프를 사용하여 재시도하거나 다른 사용 가능한 GLM 모델(예: `glm-5.1`에서 임시로 `glm-4.7` 또는 `glm-4.6`으로 전환)을 권장합니다.

## 결론

이 문서를 통해 GLM Chat Completion API를 사용하여 지표 AI의 GLM 시리즈 모델을 호출하는 방법, 기본 호출, 스트리밍 응답, 다중 대화, 시스템 프롬프트 및 도구 호출 등의 전형적인 용법을 이해하게 되었습니다. 이 문서가 API를 더 잘 연결하고 사용하는 데 도움이 되기를 바랍니다. 질문이 있으시면 언제든지 기술 지원 팀에 문의해 주십시오.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.