Can any of those flashy harnesses (Hermes, Openclaw, OpenHuman, Paperclip, Claude Code and others) run on low context?
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
Hello, running Qwen 3.8 27b Q4 K Small on a RTX A4500, 20GB VRAM, 28GB RAM I managed with Thetom Turboquant llama.cpp and some tuning, to reach an average of 31-32 tk/s ranging from 24tk/s to 43 tk/s depending also on context size. (NO VISION: 65k context, WITH VISION: 32-40k context) I tried many…
Read the full story at r/LocalLLM ↗