AINewsnow

Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork

I've been getting PrismML's Ternary Bonsai 2 27B running fast on Intel Arc. PrismML's fork only just gained basic SYCL support for its weight formats (a plain vector-dot kernel, merged 24 Sep); this goes further. Branch: https://github.com/Torchit1/llama.cpp/tree/arc-b580 (Windows zip under Release…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 10:27 · r/LocalLLM
    Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork

More stories

  1. Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? — r/LocalLLaMA
  2. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  3. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  4. Ternary bonsai 2 sur ik llama.cpp — r/LocalLLM
  5. Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation — r/AI_Agents
  6. Adaptive KV-Cache Streaming V2: Full Context MTP — r/LocalLLM
  7. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  8. Qwen 3.8 is a workhorse — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →