3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
The post describes some experiments I had while trying to desperately run deepseek-v4-flash-0731 4 bit+ quants on my machine which is supposed to support only q2 quants of the model, a or 2.xx bpw quants at best. Long story short , I wanted to have my tgs in the high twenties and my prompt processi…
Read the full story at r/LocalLLaMA ↗