What is the best AI model and quantization to run the Hermes agent comfortably on 16GB VRAM?
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
I want to try running scheduled tasks using local AI models in Hermes. Which AI model is best to use in the Hermes agent? How do you handle context, and what quantization techniques should be used to run it comfortably on 16GB VRAM?
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-10 15:08 · r/LocalLLM
What is the best AI model and quantization to run the Hermes agent comfortably on 16GB VRAM?