REDACTED1/llm-chat-demo
0
Qwen2.5-1.5B-Instruct — in your browser, free
A chat demo powered by transformers.js running the model entirely in your browser via WebGPU — nothing is sent to a server.
- Model: onnx-community/Qwen2.5-1.5B-Instruct (q4f16, ~1.2 GB, cached after first load)
- Requires a WebGPU-capable browser (Chrome/Edge 113+, recent Firefox/Safari)
- Falls back to WASM (slower) when WebGPU is unavailable
First load downloads ~1.2 GB of weights, then everything runs locally.
