AINewsnow

Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532.

Edit: Full" means the regular (non-ternary) Qwen3.8-27B, as the official NInfer Q4/Q5 groupwise quant (~4.5 bits/weight), not BF16. On my last post I showed Ternary Bonsai 2 27B running at 500+ tok/s in NInfer, a from-scratch C++/CUDA engine, on one RTX 4090 under native Windows (no WSL, no Docker)…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-28 12:43 · r/LocalLLM
    Arc Pro B70 + Qwen3.8 27B Q4_K_M: llama.cpp SYCL vs Vulkan llama-bench results
  2. 2026-09-26 14:22 · r/LocalLLM
    Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532.

More stories

  1. We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. — r/LocalLLM
  2. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  3. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  4. Adaptive KV-Cache Streaming V2: Full Context MTP — r/LocalLLM
  5. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  6. Qwen 3.8 is a workhorse — r/LocalLLaMA
  7. 95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090 — r/LocalLLaMA
  8. FIXED: HTTP 400: Failed to load model "[Specific_Model_Name_In_Use_HERE]". Error: Engine protocol runtime llama-server for [your_chat_session_number_HERE] exited before becoming healthy. exitCode=1, signal=null — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →