AINewsnow

Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp

This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab. LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weight…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-06 06:41 · r/LocalLLaMA
    Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp

More stories

  1. Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM — r/ArtificialInteligence
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Photon Announces $4.5M Seed Round to Help Developers Build AI Agents for iMessage and WhatsApp — AI Insider
  4. Overclocking DDR5 For Faster MoE Prefill and Decode — r/LocalLLaMA
  5. llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more — r/LocalLLaMA
  6. Llama.cpp + WebGPU = agants.html — r/AI_Agents
  7. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA
  8. WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →