Help me understand gguf size/ctx size
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
Let's say I have 2x 16Gb GPUs and I want to run Qwen3.8 27B. Monitor is ran by the integrated GPU so both 16Gb GPUs are almost fully free. I load the UD-Q4_K_S on one card at 15.4Gb. I then load the context on the other card? Would that be the most efficient way? Or should I aim for higher quants t…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-04 21:57 · r/LocalLLaMA
Help me understand gguf size/ctx size