Skip to main content
GLM (General Language Model) is a new generation of large language model series launched by Zhipu AI (Zhipu AI / Z.ai), which has strong capabilities in understanding and generating both Chinese and English. It performs excellently in tasks such as Chinese scenarios, code generation, reasoning, and multi-turn dialogue. The new generation models like GLM-5.3, GLM-5.2, and GLM-4.7 have made significant optimizations in long context, tool invocation, and code tasks, and can be widely applied in scenarios such as intelligent Q&A, content creation, code assistance, and customer service robots. This document mainly introduces the usage process of the GLM Chat Completion API, which allows you to easily call the GLM series models through a unified OpenAI compatible interface.

Application Process

To use the GLM Chat Completion API, first go to the Ace Data Cloud Console to obtain your API Token for backup. If you are not logged in or registered, you will be automatically redirected to the login page to invite you to register and log in. After completing this, you will be automatically returned to the current page. One API Token can call all services on the platform, no need to apply separately for each service. The first application will grant a free quota for a free experience; when the quota is insufficient, you can recharge the general balance in the console.
📘 Complete documentation: GLM Chat Completion API →

Basic Usage

The request address for the GLM Chat Completion API is https://api.acedata.cloud/glm/chat/completions, using Bearer Token for authentication, and the request body is compatible with the OpenAI Chat Completions protocol. When using this interface for the first time, we need to fill in at least three contents:
  • authorization: Select Bearer Token directly from the dropdown list.
  • model: Choose the GLM model to call, currently supported models include:
    • glm-5.3: The latest flagship model, supports 1M context and a maximum output of 128K, suitable for complex reasoning, code, and Agent tasks. Reasoning is always enabled, and you can choose low, high, or max for reasoning_effort.
    • glm-5.2: The previous generation flagship model, with strong overall capabilities.
    • glm-5.1: A mature flagship model, suitable for general complex tasks.
    • glm-4.7: Performs excellently in reasoning, tool invocation, and code tasks.
    • glm-4.6: A general dialogue model, balancing effect and cost.
    • glm-3-turbo: A classic dialogue model, suitable for general text generation tasks.
  • messages: An array of prompts, each message contains role and content, with role supporting three roles: user, assistant, system.
Common optional parameters:
  • max_tokens: Limits the maximum number of tokens in a single reply.
  • temperature: Randomness of generation, between 0-2, with larger values being more divergent.
  • top_p: Nucleus sampling parameter, controlling the cumulative probability threshold of candidate tokens.
  • n: How many candidate replies to generate at once.
  • stream: Whether to enable streaming response, default is false.
  • stop: Custom stop sequence.
Here is a simple Python call example:
After the call, we find the returned result as follows:
The main fields of the returned result are explained as follows:
  • id: The unique ID of this dialogue task.
  • created: The creation time of this dialogue task (Unix timestamp, in seconds).
  • model: The name of the GLM model actually called.
  • choices: The list of replies generated by the model. choices[i].message.content is the specific text of the model’s reply, and finish_reason indicates the reason for ending (e.g., stop, length, tool_calls, content_filter, etc.).
  • usage: Token usage statistics for this request, including prompt_tokens, completion_tokens, and total_tokens.

Streaming Response

This interface supports streaming responses (Server-Sent Events), which is very useful for web integration, allowing the webpage to achieve a word-by-word display effect. To return responses in a streaming manner, set the stream parameter in the request body to true. Python sample call code:
The output effect is as follows (excerpt):
You can see that there are many data in the response, each containing an incremental fragment. choices[i].delta.content is the newly added text fragment for the current chunk, and you can concatenate these fragments to form a complete reply. When the content of data is [DONE], it indicates the end of the streaming response. The last chunk with usage will summarize the token usage for this request. JavaScript (Node.js) example:
Java example code:
Other languages can be rewritten similarly; the principle is the same.

Multi-turn Dialogue

If you want to implement multi-turn dialogue functionality, you need to sequentially place the historical dialogue into the messages array, maintaining the alternating order of user and assistant. Python example call code:
By uploading multiple questions, you can easily achieve multi-turn dialogue and receive responses like the following:
As you can see, the information contained in choices is consistent with basic usage, and the model provides replies based on the complete dialogue history, thus supporting multi-turn contextual interaction.

System Prompt

You can add a message with a role of system at the beginning of messages to constrain the model’s role, style, or behavior:

Function Calling

The GLM model supports OpenAI-compatible Function Calling, which can declare callable functions through the tools parameter. The model will return structured function call information in choices[i].message.tool_calls when needed.
If the model decides to call a tool, the finish_reason in the returned result will change to tool_calls, and the function name and parameters in JSON string format will be provided in message.tool_calls. You can execute that function and return the result as a message with role of tool back to the model, thus completing the full tool calling loop.

Model Selection Recommendations

When api_error is returned and the message is Service is temporarily unavailable, please retry later., it usually indicates that the upstream GLM service is temporarily unavailable. It is recommended to retry with exponential backoff or switch to other available GLM models (for example, temporarily switch from glm-5.1 to glm-4.7 or glm-4.6).

Conclusion

Through this document, you have learned how to use the GLM Chat Completion API to call Zhipu AI’s GLM series models, including basic calls, streaming responses, multi-turn dialogues, system prompts, and tool invocations. We hope this document helps you better integrate and use this API. If you have any questions, please feel free to contact our technical support team.