AINewsnow

Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context, ~80-90 t/s code, 250+ t/s edits, 44 t/s at 115K

I've been getting PrismML's Ternary Bonsai 2 27B running fast on Intel Arc. PrismML's fork only recently gained basic SYCL support for its weight formats (a plain vector-dot kernel, merged 24 Sep); this goes further. Branch: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (Windows zip under [Re…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-06 20:35 · r/LocalLLM
    Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context, ~80-90 t/s code, 250+ t/s edits, 44 t/s at 115K

More stories

  1. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Local AI ecosystem overview — r/LocalLLaMA
  4. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp — r/LocalLLaMA
  6. Overclocking DDR5 For Faster MoE Prefill and Decode — r/LocalLLaMA
  7. Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected) — r/LocalLLM
  8. llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →