Run SmolLM2 360M in your browser
The lightest chat model here. Answers instantly on almost any GPU; keep questions simple.
Use SmolLM2 360M in the browser All models Downloads once · then works offline
- Family
- SmolLM
- Published by
- Hugging Face
- Parameters
- 360M
- Quantisation
- q4f16_1 · 4-bit
- Download
- 194 MB · 7 files
- GPU memory
- 376 MB
- Context window
- 4k tokens
- Reasoning
- —
- Licence
- Apache 2.0
- Base model
- HuggingFaceTB/SmolLM2-360M-Instruct
- Good at
- Chat
How big is the SmolLM2 360M download?
194 MB for the q4f16_1 build, in 7 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run SmolLM2 360M?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 376 MB of GPU memory. The tool checks your GPU, memory and free space and tells you whether SmolLM2 360M fits before anything downloads.
Is anything I type sent to a server?
No. SmolLM2 360M runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
Which build of SmolLM2 360M should I choose?
Start with q4f16_1 — it is the smallest download and the fastest to run. The f32 builds are larger and use more memory, but they run on GPUs without 16-bit shader support. Builds marked 1k hold a shorter conversation in exchange for less memory.
How long a conversation can SmolLM2 360M hold?
Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
What licence is SmolLM2 360M released under?
The weights are Apache 2.0, taken from the model card of HuggingFaceTB/SmolLM2-360M-Instruct. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.