AINewsnow

Qwen 3.8 27B on dual RTX 5060 Ti - 130 tok/s - full 256k context

I have been experimenting with Qwen 3.8 27B NVFP4 on my dual RTX 5060 Ti 16GB rig, for a total of 32GB VRAM. I'm running stock vLLM 0.30.0 with no patches, on headless Ubuntu, so no desktop is using any VRAM. This shows that it's possible to run Qwen 3.8 27B without an RTX 5090. My GPUs are power l…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-22 14:15 · r/LocalLLM
    Qwen 3.8 27B on dual RTX 5060 Ti - 130 tok/s - full 256k context

More stories

  1. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  2. Help? — r/GeminiAI
  3. Qwen-Image-2.1 — r/StableDiffusion
  4. GPT image 2.5 vs Nano Banana pro vs Nano Banana 2 vs Qwen image 3 vs Seedream 5.0 pro — r/GeminiAI
  5. XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face — r/LocalLLaMA
  6. yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash — r/LocalLLaMA
  7. I built a small local studio to try Qwen-Image-2.1 on my Mac — sharing in case you want to test it too — r/StableDiffusion
  8. Alibaba's Qwen-Audio 3.1 Slashes Voice API Prices by up to 95% — AlphaSignal

Get the daily brief of stories like this at 6:30 every morning →