k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support
Finally finished my fork: https://github.com/IceFog72/ik_llama.cpp Nothing else I wanted to add/try currently works In short, it now has: Bandwidth-adaptive CPU/GPU execution Shared expert-residency ideas from https://arxiv.org/pdf/2608.16157 integrated into this fork's profiler, bounded cache, and…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 23:33 · r/LocalLLaMA
k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support
More stories
- Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
- Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
- Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
- OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
- Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
- Nvidia-backed Reflection AI unveils its first open model, Beam. Could it be America’s best chance to compete with China? — Fortune AI
- Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
- can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →