Run TinyLlama 1.1B Chat v0.4 in your browser
TinyLlama 1.1B Chat v0.4 is a 1.1B TinyLlama model. In the browser it downloads 590 MB and needs about 697 MB of GPU memory.
Use TinyLlama 1.1B Chat v0.4 in the browser All models Downloads once · then works offline
- Family
- TinyLlama
- Published by
- TinyLlama
- Parameters
- 1.1B
- Quantisation
- q4f16_1 · 4-bit
- Download
- 590 MB · 24 files
- GPU memory
- 697 MB
- Context window
- 2k tokens
- Reasoning
- —
- Licence
- Apache 2.0
- Base model
- TinyLlama/TinyLlama-1.1B-Chat-v0.4
- Good at
- Chat
How big is the TinyLlama 1.1B Chat v0.4 download?
590 MB for the q4f16_1 build, in 24 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run TinyLlama 1.1B Chat v0.4?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 697 MB of GPU memory. The tool checks your GPU, memory and free space and tells you whether TinyLlama 1.1B Chat v0.4 fits before anything downloads.
Is anything I type sent to a server?
No. TinyLlama 1.1B Chat v0.4 runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
Which build of TinyLlama 1.1B Chat v0.4 should I choose?
Start with q4f16_1 — it is the smallest download and the fastest to run. The f32 builds are larger and use more memory, but they run on GPUs without 16-bit shader support. Builds marked 1k hold a shorter conversation in exchange for less memory.
How long a conversation can TinyLlama 1.1B Chat v0.4 hold?
Its context window is 2k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
What licence is TinyLlama 1.1B Chat v0.4 released under?
The weights are Apache 2.0, taken from the model card of TinyLlama/TinyLlama-1.1B-Chat-v0.4. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.