AINewsnow

Dual B60 24GB Performance

Hi, does anyone have benchmark numbers for a dual Intel Arc Pro B60 24GB Setup? I would really be interested in: - llama SYSCTL/Vulkan - Qwen 3.8 27B Q8 - KV-Cache Q8 - Context 64 and 128bit If possible compared to the same model but at Q6_K. And if someone could run the exact same benchmark on Dua…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-22 21:06 · r/LocalLLM
    Dual B60 24GB Performance

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  3. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  4. My contribution to the local AI community: 9 abliterated models, 99 GGUF quantizations in progress — r/huggingface
  5. Is llama.cpp meant to be slow at long context, even when you aren't using that context? — r/LocalLLaMA
  6. How I structured 50k synthetic ICD-10 QA pairs for local LLM fine-tuning — r/deeplearning
  7. Who's getting above 50 tok/s on AMD 9070, R9700 GPUs? — r/LocalLLM
  8. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →