Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass
This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.
Benchmarks of the same model on the same GPU across three serving stacks, then an FP8 pass on the winner. All numbers measured on our own hardware last week. Raw CSVs, the environment manifest and a one-command reproduction script exist for every figure; the script was re-run end to end after the r…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-08-25 16:56 · DEV Community — Machine Learning
Qwen3-8B on workstation Blackwell: vLLM vs SGLang vs llama.cpp, plus an FP8 pass