Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now

GLM 5.3

GLM-5.3 is the flagship open-weights model in Z.ai's GLM-5 series: a 753-billion-parameter Mixture-of-Experts model with 40B active parameters per token and a one-million-token context window, built for complex coding and long-horizon agentic tasks.

Model

GLM-5.3 uses the same base model as GLM-5.2; all of its gains come from post-training. It pairs a sparse Mixture-of-Experts with DeepSeek Sparse Attention and multi-head latent attention to reduce attention computation and KV-cache size, and reuses a single attention indexer across every four layers, cutting per-token FLOPs by 2.9x at a one-million-token context. Post-training focused on complex coding and long-horizon agentic work: Z.ai reports a 50% improvement over GLM-5.2 on its internal code bench and state-of-the-art vulnerability discovery on CyberGym. The thinking budget is adjustable via the reasoning_effort parameter (low, high, or max). Full details are in Z.ai's announcement.

Shared Endpoint

On a Shared Endpoint, you pay per token. The endpoint is OpenAI-compatible and already live: point your existing SDK at it and start sending requests.

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute