Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Gemma 4 31B IT

Gemma 4 31B is the flagship dense model in Google DeepMind's Gemma 4 family: 30.7 billion parameters, native text and image input, a 256K-token context window, built-in configurable thinking, and function calling, released as open weights under Apache 2.0.

Model

Gemma 4 31B is a 60-layer dense transformer with a ~550M-parameter vision encoder, so it accepts images at variable aspect ratios and resolutions alongside text. Interleaved local sliding-window and global attention layers, with unified keys and values and Proportional RoPE in the global layers, keep memory in check across its 256K-token context. It ships with a switchable thinking mode for step-by-step reasoning, native function calling for agentic workflows, and out-of-the-box support for 35+ languages (pre-trained on 140+).

This is the instruction-tuned variant, released as open weights under the Apache 2.0 license. Full details are in Google's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Gemma 4 31B IT, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model google/gemma-4-31B-it

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute