AINewsnow

95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Hello everyone! A little while back I posted about LlamAmpere , a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too). Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to shar…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-28 18:35 · r/LocalLLaMA
    95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

More stories

  1. PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill — r/LocalLLM
  2. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  3. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  4. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  5. FIXED: HTTP 400: Failed to load model "[Specific_Model_Name_In_Use_HERE]". Error: Engine protocol runtime llama-server for [your_chat_session_number_HERE] exited before becoming healthy. exitCode=1, signal=null — r/LocalLLM
  6. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  7. Anyone customizing and Optimizing llama.cpp per model? — r/LocalLLaMA
  8. llama.cpp MacOS menu bar app using blobs instead of GGUF files — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →