AINewsnow

I Made a Modular Voice Agent Running Offline on an RTX 5080 Laptop GPU: Qwen3.6-35B-A3B (3-bit), Whisper + Piper, local graph memory [demo]

Everything runs local. No API calls, no cloud, no network at runtime. End-to-end (EOU → first audio): ~1.4–3.5 s, depending on audio stack and the context/task. Core model - Qwen3.6-35B-A3B (MoE, ~3B active params/token) - 3-bit quant, vision tower ablated — ~13.7 GB on disk - llama.cpp (llama-cpp-…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-26 00:22 · r/LocalLLM
    I Made a Modular Voice Agent Running Offline on an RTX 5080 Laptop GPU: Qwen3.6-35B-A3B (3-bit), Whisper + Piper, local graph memory [demo]

More stories

  1. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  2. model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  4. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  5. KV Cache Math: Why Llama 3.1 8B at 128K Context Won't Fit in 24GB — DEV Community — Machine Learning
  6. 42x Faster Prompt Lookup Drafting in llama.cpp — r/LocalLLaMA
  7. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM
  8. Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →