Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine
Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine I've been running Prism ML's Ternary Bonsai 2 27B (Qwen3.8-27B with {−1, 0, +1} weights) in NInfer, a from-scratch C++/CUDA inference engine, on a single RTX 4090…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-26 14:22 · r/LocalLLM
Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532. - 2026-09-25 22:47 · r/LocalLLM
Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine
More stories
- Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
- model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp — r/LocalLLaMA
- My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context — r/LocalLLaMA
- Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
- Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM
- Qwen3.8 27B on Intel X7 358h + B390, with pi + llama.cpp surprised by its own RAM speed — r/LocalLLM
- 2x Radeon AI PRO R9700 + Qwen3.8-27B: 31 t/s on Windows → 113 t/s on Linux/vLLM. The fix was an M.2 riser, because the chipset slot doesn't do PCIe atomics. — r/LocalLLM
- Ling Tiny 3.0 is a glimpse of the future — r/LocalLLaMA
Get the daily brief of stories like this at 6:30 every morning →