LLM Perks

GLM 4 Flash

Z.ai (Zhipu AI) · All GLM models
Free right now Mid-size New
ReasoningTool useOpen weights

GLM 4 Flash is a language model from Z.ai (Zhipu AI). It is currently free at 1 provider with up to 128K tokens of context. The model supports step-by-step reasoning, supports tool calling. Its weights are open, so it can also run locally.

Developer
Z.ai (Zhipu AI)
Context
128K tokens
Type
Text
License
Open weights
Free since
October 11, 2026

Where to use GLM 4 Flash for free

Each row is a separate free endpoint with its own limits.

ProviderContextLimitCard
UnoRouter
glm-4-flash:free
128K No-card free models use shared capacity and per-model limits; availability can change. No Open

Quick start

curl https://api.unorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $UNOROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-4-flash:free",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Example for UnoRouter. Any OpenAI-compatible client works: Cursor, Cline, OpenCode, Open WebUI.

Benchmarks

This model is not on the Artificial Analysis leaderboard yet. We do not publish our own estimates.

Similar free models

FAQ

Is GLM 4 Flash free?

Yes. As of October 11, 2026, GLM 4 Flash is free at UnoRouter. Each provider has its own limits — see the table above.

What is the context window of GLM 4 Flash?

The largest context among free routes is 128,000 tokens. Some providers cap it lower.

How do I call GLM 4 Flash via API?

Use any OpenAI-compatible client with base URL https://api.unorouter.com/v1, model ID glm-4-flash:free, and a UnoRouter key.

Is GLM 4 Flash open-weight?

Yes, the weights are published — you can download and run it locally, for example with Ollama or LM Studio.