Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork
I've been getting PrismML's Ternary Bonsai 2 27B running fast on Intel Arc. PrismML's fork only just gained basic SYCL support for its weight formats (a plain vector-dot kernel, merged 24 Sep); this goes further. Branch: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (Windows zip under Release…
Read the full story at r/LocalLLM ↗