Self-hosting Qwen3.8-Flash-Next (or a smaller alternative) for a heavy multi-agent Hermes setup on $10k of AWS credits, looking for options I've missed
I run a Kanban-driven multi-agent orchestrator on Hermes Agent: 20+ profiles, 1,000+ skills, MCP routing, and self-hosted Mem0 on pgvector. On Fireworks, I was doing roughly 9-15B tokens a month with a 90-98% cache hit rate, at about 30-60 Kanban tasks a day. I now want to move this to infrastructu…
Read the full story at r/LocalLLM ↗