AINewsnow

Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source)

We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot. Short answer: yes, if you pick the right model. Hardware and model One NVIDIA DGX Spark (128GB unified memory) Ornith-1.5-…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-06 02:59 · r/LocalLLM
    Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source)

More stories

  1. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  2. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  3. The Nvidia Shield TV Is 7 Years Old. It Just Got a $100 Price Hike — Wired AI
  4. Reflection AI releases first open model to rival China — The Hill Technology
  5. Nvidia-Backed Reflection Unveils Open AI Model, Taking on China — Bloomberg AI
  6. TagScribeR rebuilt: a free, local dataset studio with native LoRA training (AMD ROCm and NVIDIA) — r/StableDiffusion
  7. A new startup Ghost launched a personal AI computer built to run AI locally for consumers, has a RTX Pro 4000 Blackwell SFF GPU and 64GB DDR5 RAM of memory for $3499 — r/LocalLLM
  8. Scoop: A powerful new model from startup Reflection is set to shake up the AI race — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →