Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Gemma 3 1B IT

Gemma 3 1B is the smallest of the four Gemma 3 sizes Google DeepMind launched in March 2025: a 1-billion-parameter, text-only, instruction-tuned LLM with a 32K-token context window, trained on 2 trillion tokens and released as open weights under the Gemma Terms of Use for phones, laptops and your own cloud.

Model

Gemma 3 1B is a dense decoder-only transformer with roughly 1.0 billion parameters, 302M of them in embeddings and 698M non-embedding, trained by knowledge distillation on 2 trillion tokens. Unlike its 4B, 12B and 27B siblings it has no vision encoder and accepts text only, with a 32K-token input context and up to 8,192 output tokens. It uses Gemma 3's 5:1 interleaving of local sliding-window and global attention layers, grouped-query attention with QK-norm, and a 262K-entry tokenizer shared with Gemini 2.0. Pretraining data spans more than 140 languages, and the family supports function calling and structured output for agentic workflows.

This is the instruction-tuned variant, post-trained with Google's Gemma 3 recipe that improves math, chat, instruction following and multilingual ability over the pretrained checkpoint. The weights are released under the Gemma Terms of Use rather than a standard open-source license, and the Hugging Face repository requires accepting those terms before download. Full details are in Google's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Gemma 3 1B IT, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model google/gemma-3-1b-it

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute