AINewsnow

4x RTX 3090 PCIe 4.0 x16 - advice? Qwen 3.8 Next Flash?

We're upgrading our server (Threadripper Pro 5955WX, 128GB 8-channel DDR4) from two RTX 3090 to four cards. Currently we're running Qwen 3.8 27b Q8 with vLLM for a few users. I'm wondering what we should do next: Keep running Qwen 3.8 27b Q8 and enjoy the performance boost and use a bit of spare VR…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-19 10:49 · r/LocalLLaMA
    4x RTX 3090 PCIe 4.0 x16 - advice? Qwen 3.8 Next Flash?

More stories

  1. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  2. Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use — r/machinelearningnews
  3. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  4. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  5. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial
  6. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  7. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  8. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →