Gemma 4 E2B is the edge-sized member of Google DeepMind's Gemma 4 family: 2.3B effective parameters (5.1B including embeddings), native text and image input, a 128K-token context window, configurable thinking and function calling, released as open weights under Apache 2.0.
Gemma 4 E2B is a 35-layer dense transformer built for phones, edge devices and laptops. The E stands for effective parameters: Per-Layer Embeddings give every decoder layer its own small per-token embedding table, so the model has 5.1B parameters in total but only about 2.3B are active in compute. A ~150M-parameter vision encoder handles images at variable aspect ratios and resolutions with a configurable visual token budget. Hybrid local sliding-window and global attention with Proportional RoPE supports a 128K-token context, alongside a switchable thinking mode, native function calling and a system role.
This is the instruction-tuned variant, released as open weights under the Apache 2.0 license. Google recommends sampling at temperature 1.0, top-p 0.95 and top-k 64, placing images before the text in a prompt, and enabling thinking by prefixing the system prompt with the think control token. Full details are in Google's announcement and the model card.
A Dedicated Endpoint is your own deployment of Gemma 4 E2B IT, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model google/gemma-4-E2B-it