Running Qwen3.8 27B locally with DeepSeek Harness + llama.cpp on an RTX 5060 Ti 16GB
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I finally got a fully local coding-agent stack working on Windows without using Ollama or LM Studio for inference: DeepSeek Harness Web UI ↓ llama.cpp / llama-server ↓ Qwen3.8 27B Q3_K_XL GGUF ↓ RTX 5060 Ti 16GB Machine specs Windows 11 AMD Ryzen 7 5700X — 8 cores / 16 threads NVIDIA RTX 5060 Ti —…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 14:44 · r/LocalLLM
Running Qwen3.8 27B locally with DeepSeek Harness + llama.cpp on an RTX 5060 Ti 16GB