Gemma 4 26B A4B is the Mixture-of-Experts model in Google DeepMind's Gemma 4 family: 25.2 billion total parameters with only 3.8B active per token, native text and image input, a 256K-token context window, configurable thinking and function calling, released as open weights under Apache 2.0.
Gemma 4 26B A4B is a 30-layer Mixture-of-Experts transformer with 128 routed experts plus one shared expert, of which 8 are active per token, so only 3.8B of its 25.2B parameters run at inference and it decodes at close to the speed of a 4B dense model. A ~550M-parameter vision encoder accepts images at variable aspect ratios and resolutions with a configurable visual token budget. Interleaved local sliding-window and global attention, with unified keys and values and Proportional RoPE in the global layers, supports a 256K-token context. It has a switchable thinking mode, native function calling and a system role, and scores 82.6% on MMLU Pro and 77.1% on LiveCodeBench v6.
This is the instruction-tuned variant, released as open weights under the Apache 2.0 license. Google recommends sampling at temperature 1.0, top-p 0.95 and top-k 64, with thinking enabled by prefixing the system prompt with the think control token. Full details are in Google's announcement and the model card.
A Dedicated Endpoint is your own deployment of Gemma 4 26B A4B IT, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model google/gemma-4-26B-A4B-it