Z.ai
Open weights3 channelsglm-5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering
Aggregated modalities
Available from 3 channels
Each channel is compared at its lowest input-price tier; output breaks ties. A dash means /api/pricing does not provide that rate. Selecting a channel updates the endpoints and call example below.
ChannelUSD per 1M tokensInputOutputCache readCache write
jd/glm-5.2Lowest input
Input
$1.0885 USD per 1M tokens
Output
$3.8095 USD per 1M tokens
Cache read
$0.2721 USD per 1M tokens
Cache write
- USD per 1M tokens
baidu/glm-5.2Input
$1.3333 USD per 1M tokens
Output
$4.1905 USD per 1M tokens
Cache read
$0.2476 USD per 1M tokens
Cache write
- USD per 1M tokens
z-ai/glm-5.2Input
$1.40 USD per 1M tokens
Output
$4.40 USD per 1M tokens
Cache read
$0.26 USD per 1M tokens
Cache write
- USD per 1M tokens
Endpoints
POST
/v1/chat/completionsChat CompletionsCall it
Using jd/glm-5.2
1from openai import OpenAI23client = OpenAI(4 base_url="https://www.realrelay.ai/v1",5 api_key="sk-***",6)78response = client.chat.completions.create(9 model="jd/glm-5.2",10 messages=[{"role": "user", "content": "Hello"}],11)1213print(response.choices[0].message.content)1import OpenAI from "openai";23const client = new OpenAI({4 baseURL: "https://www.realrelay.ai/v1",5 apiKey: process.env.REALRELAY_API_KEY,6});78const response = await client.chat.completions.create({9 model: "jd/glm-5.2",10 messages: [{ role: "user", content: "Hello" }],11});1213console.log(response.choices[0].message.content);1curl https://www.realrelay.ai/v1/chat/completions \2 -H "Authorization: Bearer $REALRELAY_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "jd/glm-5.2",6 "messages": [{ "role": "user", "content": "Hello" }]7 }'Prices and availability as published on Oct 4, 2026.
Capabilities
MCPPrompt cachingReasoningStructured outputTool use
Input
Text
Output
Text
