Thoughts on getting a V100 32GB as a second card for local inference? Is it worth it in 2026, or buy newer?
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
Running local LLMs on a Ryzen 9 7900 + RTX 5060 Ti 16GB (small 20L build). Considering a V100 32GB to run 30B-class models fully in VRAM. Looking for real experience: 1. Still worth it? Volta's aging so that means no modern FlashAttention paths, CUDA support winding down. What tok/s are you actuall…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-16 15:18 · r/LocalLLM
Thoughts on getting a V100 32GB as a second card for local inference? Is it worth it in 2026, or buy newer?