Sanity Check: Qwen 3.8 + Deepseek Harness actively inferencing on heterogenous garbage GPUs
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
I've been working on a custom new control plane that is designed to be OS agnostic, automatically figures out what GPUs you have in your system, what models you have installed, queues up llama.cpp or vllm, dynamically selects them for duty, figures out user request concurrency, all while managing t…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-03 15:25 · r/LocalLLM
Sanity Check: Qwen 3.8 + Deepseek Harness actively inferencing on heterogenous garbage GPUs