AINewsnow

Sanity Check: Qwen 3.8 + Deepseek Harness actively inferencing on heterogenous garbage GPUs

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

I've been working on a custom new control plane that is designed to be OS agnostic, automatically figures out what GPUs you have in your system, what models you have installed, queues up llama.cpp or vllm, dynamically selects them for duty, figures out user request concurrency, all while managing t…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-03 15:25 · r/LocalLLM
    Sanity Check: Qwen 3.8 + Deepseek Harness actively inferencing on heterogenous garbage GPUs

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  4. I gave 6 different AIs the same 5 questions — r/AI_Agents
  5. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  6. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA
  7. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  8. My Version of Jev running locally, playing doom. — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →