AINewsnow

$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill

Post title is slightly misleading, I don't think you can get these for $350 each anymore but they're still pretty cheap all things considered. They're Radeon Pro V620's which are older RDNA2 enterprise cloud gaming cards with 32 GB VRAM. (Ignore the RTX 4090 on the side, it's just used for stuff li…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-08 17:12 · r/LocalLLaMA
    $2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Mistral Large 4 — Mistral AI News
  3. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  4. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  5. Sharing AI progress in mathematics — OpenAI News
  6. OpenAI Decisions API now available on AI Gateway — Vercel Blog
  7. Introducing Playground: Create and play custom games — Google AI Blog
  8. Anthropic bans ‘abusive or cruel behavior’ toward Claude — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →