Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp
This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab. LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weight…
Read the full story at r/LocalLLaMA ↗