Insane PP difference before and after with P2P hack for Nvidia RTX!
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
I have a rig with 4 x 5070 Ti. Total VRAM is 64 GB. I'm running Qwen/Qwen3.8-27B-FP8 in vLLM and am getting what I believe is good performance. I have an EPYC 7532 on an ASRock Rack ROMED8-2T motherboard, so there are enough PCIe lanes. Therefore, I've never really bothered trying to get P2P to wor…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-04 14:16 · r/LocalLLM
Insane PP difference before and after with P2P hack for Nvidia RTX!