Multiple-GPU scaling with RTX 5060 Ti / 16 GB GPsU - llama-bench
This story is from 2026-08-19. It is preserved in the archive; the latest stories are on the live feed.
I have a Windows 11 Pro box with 3 x 5060 Ti 16GB, all in PCIe 4.0 x16 slots - TR Pro 3955WX / 128GB box. I have been playing with many quants of Qwen3.8-27B using llama.cpp bench and CUDA . Using the smallest quants, I find that there is very small benefit to having the second GPU. The prompt proc…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-19 23:24 · r/LocalLLM
Multiple-GPU scaling with RTX 5060 Ti / 16 GB GPsU - llama-bench