AINewsnow

More stories

  1. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  3. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  4. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  5. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  6. GPU - Vulkan llama.cpp benchmarks sorted by price to performance — r/LocalLLaMA
  7. MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash — r/LocalLLaMA
  8. Best current R9700 inference engine? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →