GLM 5.3 Flash
GLM-5.3-Flash is the first natively multimodal model in Z.ai's GLM-5 series: a 320-billion-parameter Mixture-of-Experts model with 18B active parameters per token and a one-million-token context window, built for coding and agentic work.
Model
GLM-5.3-Flash pairs sparse attention with linear attention in a hybrid architecture that reduces attention computation and KV-cache size. Trained on a multimodal corpus, it integrates visual understanding directly into coding and agentic workflows. Full details are in Z.ai's announcement.
Shared Endpoint
On a Shared Endpoint, you pay per token. The endpoint is OpenAI-compatible and already live: point your existing SDK at it and start sending requests.
Dedicated Endpoint
DocsA Dedicated Endpoint is your own deployment of GLM 5.3 Flash, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model zai-org/GLM-5.3-Flash 
