Run NeuralHermes 2.5 Mistral 7B in your browser
NeuralHermes 2.5 Mistral 7B is a 7B Nous Research model. In the browser it downloads 3.8 GB and needs about 4.5 GB of GPU memory.
Use NeuralHermes 2.5 Mistral 7B in the browser All models Downloads once · then works offline
- Family
- Hermes
- Published by
- Nous Research
- Parameters
- 7B
- Quantisation
- q4f16_1 · 4-bit
- Download
- 3.8 GB · 106 files
- GPU memory
- 4.5 GB
- Context window
- 4k tokens
- Reasoning
- —
- Licence
- Apache 2.0
- Base model
- mlabonne/NeuralHermes-2.5-Mistral-7B
- Good at
- Chat
| # | Build | Download | |
|---|---|---|---|
| 01 | q4f16_1 | 3.8 GB | Recommended |
How big is the NeuralHermes 2.5 Mistral 7B download?
3.8 GB for the q4f16_1 build, in 106 files. It downloads once and stays in your browser cache, so every later visit starts straight away and works offline.
What does my device need to run NeuralHermes 2.5 Mistral 7B?
A browser with WebGPU — recent Chrome, Edge, Firefox or Safari — and about 4.5 GB of GPU memory. The tool checks your GPU, memory and free space and tells you whether NeuralHermes 2.5 Mistral 7B fits before anything downloads.
Is anything I type sent to a server?
No. NeuralHermes 2.5 Mistral 7B runs on your own GPU inside the browser tab. The only network traffic is the one-time download of the model files from Hugging Face; your prompts and the answers never leave the device, and the chats are stored in this browser.
How long a conversation can NeuralHermes 2.5 Mistral 7B hold?
Its context window is 4k tokens. Longer conversations keep working — the oldest turns are dropped with a notice when the conversation no longer fits.
What licence is NeuralHermes 2.5 Mistral 7B released under?
The weights are Apache 2.0, taken from the model card of mlabonne/NeuralHermes-2.5-Mistral-7B. Check the licence yourself before using it commercially.
Figures read from the WebLLM 0.2.84 catalog and the model card on 2026-09-03.