Qwen3.8-27B on a single RTX 5090 (32GB) + 64GB DDR5-6000 — looking for real t/s numbers (llama.cpp / vLLM / sglang)
Looking at this checkpoint: lued/Qwen3.8-27B-INT8-W8A16-MTP — INT8 W8A16, ~29.4GB on disk, native 262K context, BF16 MTP head for speculative decoding. It's explicitly built as an Ampere-optimized checkpoint (W8A16 Marlin path) since Ampere lacks native FP8 tensor cores. The only published benchmar…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-19 10:21 · r/LocalLLM
Qwen3.8-27B on a single RTX 5090 (32GB) + 64GB DDR5-6000 — looking for real t/s numbers (llama.cpp / vLLM / sglang)