Gemma 4 E4B is the larger edge model in Google DeepMind's Gemma 4 family: 4.5B effective parameters (8B including embeddings), native text and image input, a 128K-token context window, configurable thinking and function calling, released as open weights under Apache 2.0.
Gemma 4 E4B is a 42-layer dense transformer designed for laptops and mobile devices. Like E2B it uses Per-Layer Embeddings, so of its 8B total parameters only about 4.5B are effective in compute, which is what the E denotes. A ~150M-parameter vision encoder processes images at variable aspect ratios and resolutions with a configurable visual token budget. Hybrid local sliding-window and global attention with Proportional RoPE supports a 128K-token context, and the model offers a switchable thinking mode, native function calling and a system role. It scores 69.4% on MMLU Pro and 52.0% on LiveCodeBench v6.
This is the instruction-tuned variant, released as open weights under the Apache 2.0 license. Google recommends sampling at temperature 1.0, top-p 0.95 and top-k 64, placing images before the text in a prompt, and enabling thinking by prefixing the system prompt with the think control token. Full details are in Google's announcement and the model card.
A Dedicated Endpoint is your own deployment of Gemma 4 E4B IT, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model google/gemma-4-E4B-it