AINewsnow

Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal?

This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.

Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4 — is this config optimal? Hardware CPU: Intel Core i5-12600K RAM: 128 GB DDR4 @ 3600 MHz GPU: NVIDIA RTX 3090, 24 GB VRAM OS: Windows 11 llama.cpp: freshly compiled from today's master (build b10794, Sep 4 2026) Model Qwen3.8-Fl…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-06 19:15 · r/LocalLLM
    CPU only, 64 GB DDR5: Qwen3.8-Flash-Next UD-Q3_K_XL
  2. 2026-09-05 12:52 · r/LocalLLaMA
    Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal?

More stories

  1. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  2. Which models you run on your Nvidia v100? — r/LocalLLM
  3. Am I right in thinking llama.cpp is the only show in town for mixed (Nvidia) GPUs? — r/LocalLLM
  4. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Huang and Zuckerberg back AI safety without a slowdown: Why Big Tech is betting on market forces to police AI — Mint AI
  6. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  7. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  8. Intel releases OpenVINO 2026.4 — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →