Is it possible to run it with a combined memory setup: 16 GB VRAM + 64 GB RAM + SSD for offloading n-grams?
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
Hardware: rtx 5080 16 gb vram; 64 gb ram ddr5 6000hz; ssd with unlimited memory; ryzen 7 9800 x3d. OS: Windows 11 Software: I’d prefer llama.cpp, but it’s not a strict requirement; I’ll use whatever you suggest, as long as it works on Windows. My attempts to run it with llama.cpp: llama-server ^ -m…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-29 12:55 · r/LocalLLaMA
Is it possible to run it with a combined memory setup: 16 GB VRAM + 64 GB RAM + SSD for offloading n-grams?