AINewsnow

85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri

I've been experimenting with ways to run MoE models that don't fit comfortably in RAM, and I ended up making Overspill, a disk tier for FreeToken . The basic idea came from looking at how Colibri handles experts across disk/RAM/VRAM so I took inspiration from the general approach. Repo: https://git…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-26 17:41 · r/LocalLLaMA
    85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri

More stories

  1. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  2. Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash? — r/LocalLLaMA
  3. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  4. How do you guys give your models web browsing capabilities? — r/LocalLLaMA
  5. Which is best local AI tools that can access all like Chatgpt, DeepSeek, Grok etc — r/huggingface
  6. Gemini 3.8 flash VS DeepSeek V4.1 — r/GeminiAI
  7. Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode — r/LocalLLaMA
  8. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →