AINewsnow

Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format. Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does: 262K context (the model's full native window) ~130 tok/s decode with MTP3, 70.8% acceptance 3,591 tok/s prefill on a 9K prompt P…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-09 05:32 · r/LocalLLM
    Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s
  2. 2026-10-09 05:27 · r/LocalLLaMA
    Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

More stories

  1. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  3. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  4. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  5. GPU - Vulkan llama.cpp benchmarks sorted by price to performance — r/LocalLLaMA
  6. MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash — r/LocalLLaMA
  7. Best current R9700 inference engine? — r/LocalLLaMA
  8. struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth) — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →