[Benchmarks] Qwen3.8-27B on one DGX Spark across SGLang, vLLM and llama.cpp
This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.
I ran a 12-way Qwen3.8-27B comparison on one DGX Spark. Each engine used plain decoding plus MTP, DSpark, and DFlash2. The coding workload was a seeded 50-task HumanEval+ slice with thinking on, temperature 1.0, top-p 0.95, top-k 20, concurrency 1, and a 16,384-token completion ceiling. The numbers…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-25 06:46 · r/LocalLLM
[Benchmarks] Qwen3.8-27B on one DGX Spark across SGLang, vLLM and llama.cpp