Has anyone actually made 64k feel like 300k+ with recursive local agents?
This story is from 2026-08-23. It is preserved in the archive; the latest stories are on the live feed.
I'm running Qwen 3.8 27B locally on a single GPU. I can push the context to 131k, but I'd rather run it faster at 64k if the agent can manage context properly. What I have in mind is pretty simple: one model stays loaded the whole time main agent gets 64k when something is too big, it spawns a fres…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-23 00:54 · r/LocalLLaMA
Has anyone actually made 64k feel like 300k+ with recursive local agents?