4x 3080 20GB (modded, alibaba) + Swift 1.5 HyperQwen 3.8 27B W4A16 + DFlash2 = 140-160 t/s generation
This is a follow up of my previous post: https://www.reddit.com/r/LocalLLM/s/dz9Da0WXdW I'm writing this post to help other people who have this frankenstein card achieve better generation speeds on vLLM. MTP was never working for me, it was degrading generation speeds like crazy. After a long time…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-30 10:22 · r/LocalLLM
4x 3080 20GB (modded, alibaba) + Swift 1.5 HyperQwen 3.8 27B W4A16 + DFlash2 = 140-160 t/s generation
More stories
- Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
- Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
- Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion
- Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
- I tested Qwen Image 2.1, FLUX Klein 2, Krea 2, and Z Image in ComfyUI using the same prompts. Some interesting differences came out 👀 Full comparison here if you want to check it out! — r/comfyui
- I built Slopus, a free, open-source desktop app for generating and editing AI videos locally (Minimax H3) — r/StableDiffusion
- Sonnet 5.5 orchestrated a local Qwen 3.8 27B! — r/ClaudeAI
- Using GPT-6.1 Sol only as the planner and letting a local 27B write the code cut my API bill by 77% — r/ChatGPT
Get the daily brief of stories like this at 6:30 every morning →