1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web page. No install, no…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-09 16:27 · r/LocalLLM
1-bit 27B running in browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop - 2026-09-09 13:49 · r/LocalLLaMA
1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)