Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

GLM 4.7

GLM-4.7 is Z.ai's open-weights coding partner for agentic workflows: a 358-billion-parameter Mixture-of-Experts with 32B active per token, a 200K-token context window, and interleaved, preserved, and turn-level thinking modes, released under the MIT license.

Model

GLM-4.7 is a 92-layer Mixture-of-Experts transformer with 160 routed experts plus one shared expert, activating 8 experts per token for 32B active parameters out of 358B total, and it supports a 200K-token context window. Z.ai positions it as a coding partner: compared with GLM-4.6 it improves multilingual agentic coding and terminal tasks, produces cleaner web pages and slides, uses tools more reliably, and delivers stronger mathematical reasoning. It extends the Interleaved Thinking introduced in GLM-4.5, thinking before every response and tool call, and adds Preserved Thinking, which keeps reasoning blocks across turns in coding-agent sessions, and Turn-level Thinking, which lets you enable or disable reasoning per request.

GLM-4.7 is released under the MIT license. The model is text-in, text-out, with a Preserved Thinking mode for multi-turn agentic tasks. Z.ai's API documentation lists a 200K context length and up to 128K output tokens, plus function calling, structured output, and context caching. Full details are in Z.ai's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of GLM 4.7, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model zai-org/GLM-4.7

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute