How do i make an llm do a really big prompt in 1 shot on low end hardware. I'd rather it run for 8 hrs than say continue every 3 seconds
Every harness i use for local ai, there's always a max output token limit so i always have to tell it to continue. What do I do if i have a huge prompt and would rather let the ai work for 8 hours straight? Forgive me I am a beginner and also a moron. Also, is there any mcp or whatever that can run…
Read the full story at r/LocalLLM ↗