GLM-5.2 is Z.ai's flagship model for long-horizon coding and agentic work, quantized to NVFP4 by NVIDIA for Blackwell GPUs: 753B total parameters with 40B active per token, sparse attention with IndexShare, a solid 1M-token context, configurable reasoning effort, and an MIT license.
GLM-5.2 is a 78-layer Mixture-of-Experts model with 256 routed experts plus one shared expert and 8 experts active per token, giving 40B active parameters out of 753B total. It uses DeepSeek Sparse Attention with IndexShare, which reuses one lightweight indexer across every four sparse-attention layers to cut per-token FLOPs by 2.9x at 1M context, and an improved multi-token-prediction layer that lengthens speculative-decoding acceptance by up to 20%. Z.ai built it for long-horizon tasks: a 1M-token context that stays reliable across long coding-agent trajectories, stronger agentic coding than GLM-5.1, and reasoning effort levels of high and max to trade capability against latency and cost.
This checkpoint is NVIDIA's NVFP4 quantization of GLM-5.2, produced with NVIDIA Model Optimizer for Blackwell GPUs. Only the weights and activations of the linear operators inside the MoE experts are quantized to NVFP4; the shared expert and everything outside the expert linears are left unquantized, so the architecture and 1M-token context are unchanged from Z.ai's release. The model is text-in, text-out, and the weights are governed by the MIT license inherited from Z.ai's base model. Full details are in Z.ai's announcement and the model card.
A Dedicated Endpoint is your own deployment of GLM 5.2 NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model nvidia/GLM-5.2-NVFP4