AINewsnow

Expert expansion with llama.cpp

This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.

With the help of Glm 5.3 flash I built a custom branch of llama.cpp in order to support Expert expansion with MOE models, I've tested only on metal and It works better than my DS4 version , i need feedback from other platforms, and different models. moex-expansion

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-06 18:27 · r/LocalLLaMA
    Expert expansion with llama.cpp

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  3. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  4. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  5. Connected a local model (Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M) to GIMP via MCP tools using llama.cpp - and here's the image result from my first prompt "can you draw a picture of a flower in gimp?". Needs work. Setup follows. — r/LocalLLaMA
  6. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  7. Which models you run on your Nvidia v100? — r/LocalLLM
  8. Best Open-Source Agent Harnesses for Local LLMs in 2026 — MarkTechPost

Get the daily brief of stories like this at 6:30 every morning →