AINewsnow

Next best move for local llm?

I have a desktop setup that runs Qwen 3.8 27B but slow with ~100K tokens context, and runs Qwen 3.8 Flash in Strata better but still limited to ~100k context. My use case is local software development and tinkering, as well as the occasional image/video génération in ComyUI. I feel a bit limited in…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-08 14:17 · r/LocalLLM
    Next best move for local llm?

More stories

  1. Strata: Qwen 3.8 Flash next in loop — r/LocalLLM
  2. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature — r/LocalLLaMA
  4. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
  5. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  6. RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to — r/LocalLLM
  7. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  8. Just bought a Mac mini M5 Pro (64GB) for AI development — what’s your setup? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →