GLM-5.3 is Z.ai's flagship open-weights model for complex coding and long-horizon agentic tasks, served here as Modal's NVFP4 quantization: 753B total parameters with 40B active per token, a 1M-token context, always-on reasoning with low, high, and max effort levels, and Z.ai's GLM-5.3 License.
GLM-5.3 shares its base model with GLM-5.2, so every gain comes from scaled post-training on long-horizon task environments. Architecturally it is a 78-layer Mixture-of-Experts with 256 routed experts plus one shared expert, 8 experts active per token, 40B active parameters out of 753B total, and DeepSeek Sparse Attention with IndexShare supporting a 1M-token context. Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, open-source state of the art on Terminal Bench 3.0 and Agents' Last Exam, and emergent cyber capability on vulnerability discovery and exploitation benchmarks. Reasoning is always on, with a reasoning_effort parameter that accepts low, high, or max.
This checkpoint is Modal's NVFP4 quantization of GLM-5.3, produced from Z.ai's BF16 release with NVIDIA Model Optimizer. Only the routed-expert linear layers are quantized to NVFP4; attention, the DSA indexer, shared experts, the MTP layer, embeddings, and the output head stay in BF16. The weights are released under Z.ai's GLM-5.3 License, inherited from the base model. Full details are in Z.ai's announcement and the model card.
A Dedicated Endpoint is your own deployment of GLM 5.3, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model modal-labs/GLM-5.3-NVFP4