AINewsnow

Arc Pro B70 + Qwen3.8 27B Q4_K_M: llama.cpp SYCL vs Vulkan llama-bench results

Hi folks, I tested both llama.cpp (SYCL & Vulkan) backends against the Intel Arc Pro B70 to see which one to use for Intel Arc GPUs. Setup: Arc Pro B70 32GB, Core Ultra 265, 96GB RAM, Qwen3.8 27B Q4_K_M. Windows11, llama.cpp both are WebUI builds. llama-bench (t/s): Test SYCL Vulkan pp64 114.53 337…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-28 12:43 · r/LocalLLM
    Arc Pro B70 + Qwen3.8 27B Q4_K_M: llama.cpp SYCL vs Vulkan llama-bench results

More stories

  1. Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? — r/LocalLLaMA
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  4. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  5. Ternary bonsai 2 sur ik llama.cpp — r/LocalLLM
  6. Ternary Bonsai 2 27B on a 12 GB Intel Arc B580: 128K context all in VRAM, ~80-90 t/s code, 250+ t/s edits, 2-4x faster than the official fork — r/LocalLLM
  7. Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation — r/AI_Agents
  8. Adaptive KV-Cache Streaming V2: Full Context MTP — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →