Need support for llama.cpp with multi GPU
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
Using llama.cpp I seem to be unable to get my to GPUs working tougether correclty, so I need help somehow. Setup: 96GB RAM, one Blackwell 5000 (48GB) and one 3090 (24GB). I am trying to run the UD-Q3_K_XL quant of Deepseek4 flash which has about 120GB size. Using just the Blackwell I would put most…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-22 08:53 · r/LocalLLaMA
Need support for llama.cpp with multi GPU