AINewsnow

3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed

This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.

The post describes some experiments I had while trying to desperately run deepseek-v4-flash-0731 4 bit+ quants on my machine which is supposed to support only q2 quants of the model, a or 2.xx bpw quants at best. Long story short , I wanted to have my tgs in the high twenties and my prompt processi…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-08-22 23:34 · r/LocalLLaMA
    3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed

More stories

  1. Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash — r/LocalLLaMA
  2. DeepSeek’s Insane New Architecture — Two Minute Papers
  3. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  4. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  5. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  6. I enjoyed the daily HF papers today — r/LocalLLaMA
  7. Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading — SemiAnalysis
  8. Coming soon...... Optimized for DEEPSEEK Flash.... Though model Agnostic.... message me to test.... cem888.ai — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →