GLM 4 Flash
GLM 4 Flash is a language model from Z.ai (Zhipu AI). It is currently free at 1 provider with up to 128K tokens of context. The model supports step-by-step reasoning, supports tool calling. Its weights are open, so it can also run locally.
- Developer
- Z.ai (Zhipu AI)
- Context
- 128K tokens
- Type
- Text
- License
- Open weights
- Free since
- October 11, 2026
Where to use GLM 4 Flash for free
Each row is a separate free endpoint with its own limits.
Quick start
curl https://api.unorouter.com/v1/chat/completions \
-H "Authorization: Bearer $UNOROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4-flash:free",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.unorouter.com/v1",
api_key="YOUR_UNOROUTER_API_KEY",
)
reply = client.chat.completions.create(
model="glm-4-flash:free",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.unorouter.com/v1",
apiKey: process.env.UNOROUTER_API_KEY,
});
const reply = await client.chat.completions.create({
model: "glm-4-flash:free",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(reply.choices[0].message.content);Example for UnoRouter. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.
Benchmarks
This model is not on the Artificial Analysis leaderboard yet. We do not publish our own estimates.
Similar free models
FAQ
Is GLM 4 Flash free?
Yes. As of October 11, 2026, GLM 4 Flash is free at UnoRouter. Each provider has its own limits — see the table above.
What is the context window of GLM 4 Flash?
The largest context among free routes is 128,000 tokens. Some providers cap it lower.
How do I call GLM 4 Flash via API?
Use any OpenAI-compatible client with base URL https://api.unorouter.com/v1, model ID glm-4-flash:free, and a UnoRouter key.
Is GLM 4 Flash open-weight?
Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.