Application Process
To use the GLM Chat Completion API, first go to the Ace Data Cloud Console to obtain your API Token for backup.
If you are not logged in or registered, you will be automatically redirected to the login page to invite you to register and log in. After completing this, you will be automatically returned to the current page.
One API Token can call all services on the platform, no need to apply separately for each service. The first application will grant a free quota for a free experience; when the quota is insufficient, you can recharge the general balance in the console.
📘 Complete documentation: GLM Chat Completion API →
Basic Usage
The request address for the GLM Chat Completion API ishttps://api.acedata.cloud/glm/chat/completions, using Bearer Token for authentication, and the request body is compatible with the OpenAI Chat Completions protocol.
When using this interface for the first time, we need to fill in at least three contents:
authorization: Select Bearer Token directly from the dropdown list.model: Choose the GLM model to call, currently supported models include:glm-5.3: The latest flagship model, supports 1M context and a maximum output of 128K, suitable for complex reasoning, code, and Agent tasks. Reasoning is always enabled, and you can chooselow,high, ormaxforreasoning_effort.glm-5.2: The previous generation flagship model, with strong overall capabilities.glm-5.1: A mature flagship model, suitable for general complex tasks.glm-4.7: Performs excellently in reasoning, tool invocation, and code tasks.glm-4.6: A general dialogue model, balancing effect and cost.glm-3-turbo: A classic dialogue model, suitable for general text generation tasks.
messages: An array of prompts, each message containsroleandcontent, withrolesupporting three roles:user,assistant,system.
max_tokens: Limits the maximum number of tokens in a single reply.temperature: Randomness of generation, between 0-2, with larger values being more divergent.top_p: Nucleus sampling parameter, controlling the cumulative probability threshold of candidate tokens.n: How many candidate replies to generate at once.stream: Whether to enable streaming response, default isfalse.stop: Custom stop sequence.
id: The unique ID of this dialogue task.created: The creation time of this dialogue task (Unix timestamp, in seconds).model: The name of the GLM model actually called.choices: The list of replies generated by the model.choices[i].message.contentis the specific text of the model’s reply, andfinish_reasonindicates the reason for ending (e.g.,stop,length,tool_calls,content_filter, etc.).usage: Token usage statistics for this request, includingprompt_tokens,completion_tokens, andtotal_tokens.
Streaming Response
This interface supports streaming responses (Server-Sent Events), which is very useful for web integration, allowing the webpage to achieve a word-by-word display effect. To return responses in a streaming manner, set thestream parameter in the request body to true.
Python sample call code:
data in the response, each containing an incremental fragment. choices[i].delta.content is the newly added text fragment for the current chunk, and you can concatenate these fragments to form a complete reply. When the content of data is [DONE], it indicates the end of the streaming response. The last chunk with usage will summarize the token usage for this request.
JavaScript (Node.js) example:
Multi-turn Dialogue
If you want to implement multi-turn dialogue functionality, you need to sequentially place the historical dialogue into themessages array, maintaining the alternating order of user and assistant.
Python example call code:
choices is consistent with basic usage, and the model provides replies based on the complete dialogue history, thus supporting multi-turn contextual interaction.
System Prompt
You can add a message with arole of system at the beginning of messages to constrain the model’s role, style, or behavior:
Function Calling
The GLM model supports OpenAI-compatible Function Calling, which can declare callable functions through thetools parameter. The model will return structured function call information in choices[i].message.tool_calls when needed.
finish_reason in the returned result will change to tool_calls, and the function name and parameters in JSON string format will be provided in message.tool_calls. You can execute that function and return the result as a message with role of tool back to the model, thus completing the full tool calling loop.
Model Selection Recommendations
api_error is returned and the message is Service is temporarily unavailable, please retry later., it usually indicates that the upstream GLM service is temporarily unavailable. It is recommended to retry with exponential backoff or switch to other available GLM models (for example, temporarily switch from glm-5.1 to glm-4.7 or glm-4.6).

