Got an old slow low vram GPU laying around? Might be worth it to use for Just Vision mmproj llama.cpp
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
For many, Vram is precious, I see many people recommend using --no-mmproj-offload to save gpu vram but it is painfully slow. Especially if you are using it with agentic coding. If possible, add that secondary gpu just for mmproj with --mmdev CUDA1(your gpu). It will be a magnitude faster than --no-…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-12 01:18 · r/LocalLLaMA
Got an old slow low vram GPU laying around? Might be worth it to use for Just Vision mmproj llama.cpp