Run Qwen3 8B in your browser
Large Qwen3 with a reasoning mode; the most capable Qwen3 that runs in a browser.
Use Qwen3 8B in the browser All models Downloads once · then works offline
- Family
- Qwen
- Published by
- Alibaba Cloud
- Parameters
- 8B
- Quantisation
- q4f16_1 · 4-bit
- Download
- 4.3 GB · 113 files
- GPU memory
- 5.6 GB
- Context window
- 4k tokens
- Reasoning
- Optional
- Licence
- Apache 2.0
- Base model
- Qwen/Qwen3-8B
- Weights
- mlc-ai/Qwen3-8B-q4f16_1-MLC
- Good at
- Chat · Reasoning · Multilingual
How big is the Qwen3 8B download?
4.3 GB for the q4f16_1 build, in 113 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run Qwen3 8B?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 5.6 GB of GPU memory. The tool checks your GPU, memory and free space and tells you whether Qwen3 8B fits before anything downloads.
Is anything I type sent to a server?
No. Qwen3 8B runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
Which build of Qwen3 8B should I choose?
Start with q4f16_1 — it is the smallest download and the fastest to run. The f32 builds are larger and use more memory, but they run on GPUs without 16-bit shader support. Builds marked 1k hold a shorter conversation in exchange for less memory.
How long a conversation can Qwen3 8B hold?
Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
Does Qwen3 8B show its reasoning?
Qwen3 8B can think before it answers. Turn reasoning on in the chat settings; the thinking streams into a collapsible block above the answer and is never sent back as part of the conversation.
What licence is Qwen3 8B released under?
The weights are Apache 2.0, taken from the model card of Qwen/Qwen3-8B. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.