Skip to the tool workspace
Browser 106
no file limit
Run AI in the browser 65 models · 13 families

Local AI models you can run in your browser

Every model below runs on your own GPU through WebGPU — it downloads once, then answers offline, and nothing you type leaves the device. Pick one by the size your machine can hold.

engineWebLLM 0.2.84 smallest194 MB networkthe model downloads once

Open the chat It proposes a model that fits your device

Recommended Name · note · download · size
01 SmolLM2 360M The lightest chat model here. Answers instantly on almost any GPU; keep questions simple. 194 MB Tiny 02 Qwen3 0.6B Small and quick, with an optional reasoning mode for harder questions. 320 MB Tiny 03 Qwen3.5 0.8B The newest small Qwen. A fast first model with an optional reasoning mode. 404 MB Tiny 04 Gemma 3 1B Google’s compact instruction-tuned model; needs very little GPU memory. 537 MB Tiny 05 Llama 3.2 1B Meta’s small instruct model. Low memory use and dependable for everyday questions. 663 MB Small 06 Qwen2.5 Coder 1.5B Tuned for code: completions, explanations and small refactors. 828 MB Small 07 Qwen3 1.7B A clear step up from the 0.6B in quality, still comfortable on integrated GPUs. 923 MB Small 08 Qwen3.5 2B The newest Qwen at this size. Good general chat with an optional reasoning mode. 1010 MB Small 09 Gemma 2 2B Google’s 2B instruct model; strong for its size, needs a GPU with f16 shaders. 1.4 GB Small 10 Qwen2.5 Coder 3B Code-focused model with room for longer files and multi-step edits. 1.6 GB Medium 11 Llama 3.2 3B Meta’s 3B instruct model. A dependable all-rounder for laptops with 8 GB or more. 1.7 GB Medium 12 Ministral 3 3B Mistral’s compact instruct model, new to this catalog. 1.8 GB Medium 13 Phi-4 mini Microsoft’s small Phi-4. Strong at reasoning-style questions and maths. 2.0 GB Medium 14 Qwen3 4B Capable general model with a reasoning mode; a good fit for 8 GB machines. 2.1 GB Medium 15 Qwen3.5 4B The newest mid-size Qwen: the best balance of quality and speed in this catalog for 8 GB machines. 2.2 GB Medium 16 Mistral 7B Instruct v0.3 Mistral’s 7B instruct model. Needs a discrete GPU or a 16 GB unified-memory machine. 3.8 GB Large 17 DeepSeek R1 Distill 7B Always thinks before answering; the reasoning streams into a collapsible block. 4.0 GB Large 18 Qwen2.5 Coder 7B The strongest code model in the catalog; needs 5 GB of GPU memory. 4.0 GB Large 19 Llama 3.1 8B Meta’s 8B instruct model. Excellent general chat if your GPU has 5 GB to spare. 4.2 GB Large 20 Qwen3 8B Large Qwen3 with a reasoning mode; the most capable Qwen3 that runs in a browser. 4.3 GB Large 21 Gemma 2 9B Google’s 9B instruct model. Needs 6.5 GB of GPU memory and f16 shaders. 4.8 GB Large

Qwen

22
01 Qwen2 0.5B Instruct 0.5B · Alibaba Cloud 265 MB 0.5B 02 Qwen2.5 0.5B Instruct 0.5B · Alibaba Cloud 265 MB 0.5B 03 Qwen2.5 Coder 0.5B Instruct 0.5B · Alibaba Cloud 265 MB 0.5B 04 Qwen3 0.6B Small and quick, with an optional reasoning mode for harder questions. 320 MB 0.6B 05 Qwen3.5 0.8B The newest small Qwen. A fast first model with an optional reasoning mode. 404 MB 0.8B 06 Qwen2 1.5B Instruct 1.5B · Alibaba Cloud 828 MB 1.5B 07 Qwen2 Math 1.5B Instruct 1.5B · Alibaba Cloud 828 MB 1.5B 08 Qwen2.5 1.5B Instruct 1.5B · Alibaba Cloud 828 MB 1.5B 09 Qwen2.5 Coder 1.5B Tuned for code: completions, explanations and small refactors. 828 MB 1.5B 10 Qwen2.5 Math 1.5B Instruct 1.5B · Alibaba Cloud 828 MB 1.5B 11 Qwen3 1.7B A clear step up from the 0.6B in quality, still comfortable on integrated GPUs. 923 MB 1.7B 12 Qwen3.5 2B The newest Qwen at this size. Good general chat with an optional reasoning mode. 1010 MB 2B 13 Qwen2.5 3B Instruct 3B · Alibaba Cloud 1.6 GB 3B 14 Qwen2.5 Coder 3B Code-focused model with room for longer files and multi-step edits. 1.6 GB 3B 15 Qwen3 4B Capable general model with a reasoning mode; a good fit for 8 GB machines. 2.1 GB 4B 16 Qwen3.5 4B The newest mid-size Qwen: the best balance of quality and speed in this catalog for 8 GB machines. 2.2 GB 4B 17 Qwen2 7B Instruct 7B · Alibaba Cloud 4.0 GB 7B 18 Qwen2 Math 7B Instruct 7B · Alibaba Cloud 4.0 GB 7B 19 Qwen2.5 7B Instruct 7B · Alibaba Cloud 4.0 GB 7B 20 Qwen2.5 Coder 7B The strongest code model in the catalog; needs 5 GB of GPU memory. 4.0 GB 7B 21 Qwen3 8B Large Qwen3 with a reasoning mode; the most capable Qwen3 that runs in a browser. 4.3 GB 8B 22 Qwen3.5 9B 9B · Alibaba Cloud 4.7 GB 9B

Llama

08

Gemma

05

Mistral

05

Phi

06

SmolLM

03

DeepSeek R1

02

Hermes

07

OLMo

02

TinyLlama

02

RedPajama

01

StableLM

01

WizardMath

01

Sizes and licences read from the WebLLM 0.2.84 catalog on 2026-09-03.

/ai/browser-based-ai-chat
1 tool selected
Runs locally·The model downloads once·Your input never leaves
Out0 B
Ready