AINewsnow

Need maybe say "Use llama.cpp"

So I tried that miracle engine everyone is talking about. Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation: Can you help with the following problem? So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights model…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-04 09:37 · r/LocalLLaMA
    Need maybe say "Use llama.cpp"

More stories

  1. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. Kimi K3: A Claude clone or something else? — CoreWeave Blog
  4. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant — r/LocalLLaMA
  6. FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
  7. Built a gateway so you can call DeepSeek, Qwen, Kimi, GLM, MiniMax with one key — USD billing, OpenAI-compatible — r/LocalLLM
  8. I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →