Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?
There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. ๐ตโ๐ซ Adjusting context will increase/decrease usage on both side. --tensor-split 499,501 = GPU1 12.5 GB, GPU2 15.โฆ
Read the full story at r/LocalLLaMA โ
Timeline ยท 1 report
- 2026-09-27 20:18 ยท r/LocalLLaMA
Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?