AINewsnow

Considering the recent advancement in models like Qwen 3.8 and inference like Strata, Free Tokens. How much of a gap is there between slow 128GB VRAM vs 16GB(5080)+96GB RAM

My question is what is correct upgrade path to a 5080+96GB RAM RTX PRO 48GB x 1 (8K $) DGX Spark x 1 (5.5K $) 128GB Mac M5 (6K $) Use case is local LLM that is good enough and Minimax H3.

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-05 18:24 · r/LocalLLM
    Considering the recent advancement in models like Qwen 3.8 and inference like Strata, Free Tokens. How much of a gap is there between slow 128GB VRAM vs 16GB(5080)+96GB RAM

More stories

  1. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
  2. Built a gateway so you can call DeepSeek, Qwen, Kimi, GLM, MiniMax with one key — USD billing, OpenAI-compatible — r/LocalLLM
  3. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  4. The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
  5. Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen — r/LocalLLaMA
  6. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  7. My frontier class agent fact-checks my local AI before I grade it. How do you grade your Agents and LLMs? — r/AI_Agents
  8. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →