Run Llama 3.1 70B Instruct in your browser
Llama 3.1 70B Instruct is a 70B Meta model. In the browser it downloads 30 GB and needs about 30 GB of GPU memory.
Use Llama 3.1 70B Instruct in the browser All models Downloads once · then works offline
- Family
- Llama
- Published by
- Meta
- Parameters
- 70B
- Quantisation
- q3f16_1 · 3-bit
- Download
- 30 GB · 456 files
- GPU memory
- 30 GB
- Context window
- 4k tokens
- Reasoning
- —
- Licence
- Llama 3.1 Community License
- Base model
- meta-llama/Meta-Llama-3.1-70B-Instruct
- Good at
- Chat · Multilingual
| # | Build | Download | |
|---|---|---|---|
| 01 | q3f16_1 | 30 GB | Recommended |
How big is the Llama 3.1 70B Instruct download?
30 GB for the q3f16_1 build, in 456 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run Llama 3.1 70B Instruct?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 30 GB of GPU memory. The tool checks your GPU, memory and free space and tells you whether Llama 3.1 70B Instruct fits before anything downloads.
Is anything I type sent to a server?
No. Llama 3.1 70B Instruct runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
How long a conversation can Llama 3.1 70B Instruct hold?
Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
What licence is Llama 3.1 70B Instruct released under?
The weights are Llama 3.1 Community License, taken from the model card of meta-llama/Meta-Llama-3.1-70B-Instruct. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.