AINewsnow

Qwen3.8-27B FP8 dual GPUs

Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB) = 48 GB VRAM 32 GB DDR5 6200 RAM Arch Linux, KDE on the 5090 (takes ~1-1.5 GB VRAM) daily driver OrcaRouter Qwen3.8-27B Uncensored, block-FP8 (28.75 GiB), on vLLM 0.30.0 with pipeline parallel across both cards 4070 = rank 0 with layers 0-19 + v…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-27 11:15 · r/LocalLLaMA
    Qwen3.8-27B FP8 dual GPUs

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation — The Guardian AI
  3. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  4. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times Technology
  5. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  6. ‘Things Will Never Be Chill Again’: The Doomers Who Shaped the AI Safety Freakout — Wall Street Journal Technology
  7. The Surprising Reasons China Is Skeptical of A.I. Safety Calls — New York Times AI
  8. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →