AINewsnow

Imma just say it, Strata absolutely clowned llama.cpp

So, I've been begging llama.cpp to do MoE caching for about a year, and watching them d ck around with 1% here and 2% improvements there instead... Until Strata ( https://github.com/Niko1221/Strata ) clowned llama with 5-10x prefill and 3-4x decode in about two weeks. There were numerous llama PRs…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-03 14:05 · r/LocalLLaMA
    Imma just say it, Strata absolutely clowned llama.cpp

More stories

  1. Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU. — r/LocalLLaMA
  2. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  3. How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks? — r/LocalLLaMA
  4. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  5. FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
  6. Aleph Alpha lanza Kolibri, LLM alemán que activa solo 4,4% de sus parámetros — DEV Community — Machine Learning
  7. Claude and Grok built me a local monitoring setup for my AI box: two dashboards, one for the machine, one for model training — r/LocalLLM
  8. WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →