Run Gemma 3 1B in your browser
Google’s compact instruction-tuned model; needs very little GPU memory.
Use Gemma 3 1B in the browser All models Downloads once · then works offline
- Family
- Gemma
- Published by
- Parameters
- 1B
- Quantisation
- q4f16_1 · 4-bit
- Download
- 537 MB · 15 files
- GPU memory
- 711 MB
- Context window
- 4k tokens
- Reasoning
- —
- Licence
- Gemma Terms of Use
- Base model
- google/gemma-3-1b-it
- Weights
- mlc-ai/gemma3-1b-it-q4f16_1-MLC
- Good at
- Chat · Multilingual
| # | Build | Download | |
|---|---|---|---|
| 01 | q4f16_1 | 537 MB | Recommended |
How big is the Gemma 3 1B download?
537 MB for the q4f16_1 build, in 15 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run Gemma 3 1B?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 711 MB of GPU memory. The tool checks your GPU, memory and free space and tells you whether Gemma 3 1B fits before anything downloads.
Is anything I type sent to a server?
No. Gemma 3 1B runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
How long a conversation can Gemma 3 1B hold?
Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
What licence is Gemma 3 1B released under?
The weights are Gemma Terms of Use, taken from the model card of google/gemma-3-1b-it. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.