AINewsnow

Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.

Benchmarks of the same model on the same GPU across three serving stacks, then an FP8 pass on the winner. All numbers measured on our own hardware last week. Raw CSVs, the environment manifest and a one-command reproduction script exist for every figure; the script was re-run end to end after the r…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-08-25 16:56 · DEV Community — Machine Learning
    Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  3. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  4. You can use any LLM just like JEV — r/LocalLLaMA
  5. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  6. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  7. focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737) — r/LocalLLaMA
  8. I ran Opencode and PI against the same local model on 3 identical projects, same prompts, same hardware... — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →