I optimized Bonsai 27B for 8 GB VRAM and agentic work: 36 tok/s
Ternary Bonsai 2 27B in its PTQ1_0 quant is 5.95 GB of weights, so all 65 layers stay on the card with a 64k window and the KV cache at q8_0/q4_0. I measured: RTX 4060 Ti 8 GB, CUDA: 36 tok/s AMD RX 570 8 GB, Vulkan/RADV: 7 tok/s The RX 570 is why I'm posting at all. Through the fork's Vulkan path…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-22 16:10 · r/LocalLLM
I optimized Bonsai 27B for 8 GB VRAM and agentic work: 36 tok/s