Skip to the tool workspace
Browser 106
no file limit
Hugging Face Runs on your GPU

Run SmolLM2 135M Instruct in your browser

SmolLM2 135M Instruct is a 135M Hugging Face model. In the browser it downloads 257 MB and needs about 360 MB of GPU memory.

modelSmolLM2-135M-Instruct-q0f16-MLC engineWebLLM 0.2.84 networkthe model downloads once
Download257 MB
GPU memory360 MB
Context4k tokens
Parameters135M
Builds2

Use SmolLM2 135M Instruct in the browser All models Downloads once · then works offline

Open this page in a browser with JavaScript to check whether SmolLM2 135M Instruct fits this device. It needs about 360 MB of GPU memory and a 257 MB download.
Specifications Sourced from the model card
Family
SmolLM
Published by
Hugging Face
Parameters
135M
Quantisation
q0f16 · 16-bit
Download
257 MB · 8 files
GPU memory
360 MB
Context window
4k tokens
Reasoning
Licence
Apache 2.0
Good at
Chat
Builds Same weights, different precision
# Build Download GPU memory Context Note
01 q0f16 257 MB 360 MB 4k tokens 16-bit weights, f16 · needs f16 shaders Recommended
02 q0f32 513 MB 719 MB 4k tokens 32-bit weights, f32
How to run it
01 Open the chat The button above opens SmolLM2 135M Instruct in the browser AI chat, with the model already selected.
02 Download once 257 MB arrives from Hugging Face and stays in your browser cache. You can start typing while it downloads.
03 Chat privately The model runs on your GPU. Nothing you type is uploaded, and it keeps working offline.
Questions
How big is the SmolLM2 135M Instruct download?

257 MB for the q0f16 build, in 8 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.

What does my device need to run SmolLM2 135M Instruct?

A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 360 MB of GPU memory. The tool checks your GPU, memory and free space and tells you whether SmolLM2 135M Instruct fits before anything downloads.

Is anything I type sent to a server?

No. SmolLM2 135M Instruct runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.

Which build of SmolLM2 135M Instruct should I choose?

Start with q0f16 — it is the smallest download and the fastest to run. The f32 builds are larger and use more memory, but they run on GPUs without 16-bit shader support. Builds marked 1k hold a shorter conversation in exchange for less memory.

How long a conversation can SmolLM2 135M Instruct hold?

Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.

What licence is SmolLM2 135M Instruct released under?

The weights are Apache 2.0, taken from the model card of HuggingFaceTB/SmolLM2-135M-Instruct. Check the licence yourself before using it commercially.

Other models Same family, then same size

All 65 models Open the chat

Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.

/ai/browser-based-ai-chat
1 tool selected
Runs locally·The model downloads once·Your input never leaves
Out0 B
Ready