GLM-5.2 is Z.ai's flagship model for long-horizon coding and agentic work, released here as Z.ai's own FP8 checkpoint: 753B total parameters with 40B active per token, sparse attention with IndexShare, a solid 1M-token context, reasoning effort control, and an MIT license.
GLM-5.2 is Z.ai's flagship model for long-horizon tasks, and the first in the GLM series to deliver that capability on a solid 1M-token context. It is a 78-layer Mixture-of-Experts model with 256 routed experts plus one shared expert and 8 experts active per token, for 40B active parameters out of 753B total. DeepSeek Sparse Attention with IndexShare shares one lightweight indexer across every four sparse-attention layers, reducing per-token FLOPs by 2.9x at 1M context, and an improved multi-token-prediction layer raises speculative-decoding acceptance length by up to 20%. Effort levels high and max let you trade capability against speed and cost on agentic coding work.
This is Z.ai's own FP8 quantization of GLM-5.2 (FP8 E4M3 with 128 by 128 block scaling and dynamic activation scaling), with the same architecture and 1M-token context as the BF16 model. The model is text-in, text-out, and the weights are released under the MIT license, which Z.ai describes as having no regional limits. Full details are in Z.ai's announcement and the model card.
A Dedicated Endpoint is your own deployment of GLM 5.2 FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model zai-org/GLM-5.2-FP8