Behaviour of --cache-ram in router mode?
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
Hi, until now I ran llama-server with just one model at a time. Now I want to use it in router mode in order to provide different models. For my usecase I heavily rely on --cache-ram which improves speed a lot when working on big repos. I am just wondering, when setting cache-ram = 65536 in models.…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-05 08:12 · r/LocalLLaMA
Behaviour of --cache-ram in router mode?