Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs
I spent the past year independently building Slow Vale, an LLM-driven life simulation. The Chinese server now has 800+ AI residents sharing one continuously running city. This is an engineering write-up about concurrent decisions, dynamic action spaces, context caching, and the operating costs of a…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-08 09:49 · r/LocalLLaMA
Running an LLM-driven town with 800+ persistent agents: concurrency, context caching, and inference costs