AINewsnow

GPU - Vulkan llama.cpp benchmarks sorted by price to performance

This table to help anyone looking to build a budget Data Center homelab. I copied the bulk of value based, mid level, decent speed results GPUs and feed it to AI or SI and here are the recommended results. Data taken from Llama.cpp discussion thread: Performance of llama.cpp with Vulkan #10879 Ther…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-08 20:49 · r/LocalLLaMA
    GPU - Vulkan llama.cpp benchmarks sorted by price to performance

More stories

  1. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  3. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  4. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  5. MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash — r/LocalLLaMA
  6. Best current R9700 inference engine? — r/LocalLLaMA
  7. Meta's Llama 3.3 70B Now Fits on a Single 48 GB GPU — AlphaSignal
  8. struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth) — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →