AINewsnow

Self-hosting Qwen3.8-Flash-Next (or a smaller alternative) for a heavy multi-agent Hermes setup on $10k of AWS credits, looking for options I've missed

I run a Kanban-driven multi-agent orchestrator on Hermes Agent: 20+ profiles, 1,000+ skills, MCP routing, and self-hosted Mem0 on pgvector. On Fireworks, I was doing roughly 9-15B tokens a month with a 90-98% cache hit rate, at about 30-60 Kanban tasks a day. I now want to move this to infrastructu…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-04 11:24 · r/LocalLLM
    Self-hosting Qwen3.8-Flash-Next (or a smaller alternative) for a heavy multi-agent Hermes setup on $10k of AWS credits, looking for options I've missed

More stories

  1. Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows — AWS Machine Learning Blog
  2. Sweep thousands of leases for compliance using Amazon Quick and the Adjudicated Query pattern — AWS Machine Learning Blog
  3. Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI — AWS Machine Learning Blog
  4. Serve live, governed data in AI-built apps with Amazon Quick — AWS Machine Learning Blog
  5. Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors — AWS Machine Learning Blog
  6. Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS — AWS Machine Learning Blog
  7. Implementing Multi-Environment Access for Claude Platform on AWS — AWS Machine Learning Blog
  8. Simplify dashboard drill-down with the Amazon Quick Sight hierarchy filter — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →