AINewsnow

What's the best and cheapest setup for 4-6 devs?

I want to self host qwen 3.8 flash next for 4-6 devs to use on demand. What's the cheapest possible hardware setup to get good speeds that do not feel far from the typical cloud LLM speed? I assume that's around 60-80 tps to match CC or Codex? One thing to consider for your answer is I'm open to us…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-02 03:20 · r/LocalLLM
    What's the best and cheapest setup for 4-6 devs?

More stories

  1. Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
  2. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  5. Continuity update: screen replacement, image to 3D, Qwen Image 2.1, and a lot more since 3.0 — r/StableDiffusion
  6. Is anyone else running insanely long unattended loops? — r/AI_Agents
  7. Viggle turbo v0.3 for Qwen image 2.1: less grain and cleaner surfaces — r/StableDiffusion
  8. QWEN 2.1 Throttled by GPU — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →