Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532.
Edit: Full" means the regular (non-ternary) Qwen3.8-27B, as the official NInfer Q4/Q5 groupwise quant (~4.5 bits/weight), not BF16. On my last post I showed Ternary Bonsai 2 27B running at 500+ tok/s in NInfer, a from-scratch C++/CUDA engine, on one RTX 4090 under native Windows (no WSL, no Docker)…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-28 12:43 · r/LocalLLM
Arc Pro B70 + Qwen3.8 27B Q4_K_M: llama.cpp SYCL vs Vulkan llama-bench results - 2026-09-26 14:22 · r/LocalLLM
Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532.
More stories
- We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. — r/LocalLLM
- Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
- Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
- Adaptive KV-Cache Streaming V2: Full Context MTP — r/LocalLLM
- Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
- Qwen 3.8 is a workhorse — r/LocalLLaMA
- 95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090 — r/LocalLLaMA
- FIXED: HTTP 400: Failed to load model "[Specific_Model_Name_In_Use_HERE]". Error: Engine protocol runtime llama-server for [your_chat_session_number_HERE] exited before becoming healthy. exitCode=1, signal=null — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →