AINewsnow

Running Qwen3.8 27B locally with DeepSeek Harness + llama.cpp on an RTX 5060 Ti 16GB

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

I finally got a fully local coding-agent stack working on Windows without using Ollama or LM Studio for inference: DeepSeek Harness Web UI ↓ llama.cpp / llama-server ↓ Qwen3.8 27B Q3_K_XL GGUF ↓ RTX 5060 Ti 16GB Machine specs Windows 11 AMD Ryzen 7 5700X — 8 cores / 16 threads NVIDIA RTX 5060 Ti —…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-11 14:44 · r/LocalLLM
    Running Qwen3.8 27B locally with DeepSeek Harness + llama.cpp on an RTX 5060 Ti 16GB

More stories

  1. Which models you run on your Nvidia v100? — r/LocalLLM
  2. OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news — ThursdAI
  3. Bolt Adds DeepSeek V4.1 Flash at 10x Cheaper Than V4 Pro — AlphaSignal
  4. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  5. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  6. DeepSeek’s Insane New Architecture — Two Minute Papers
  7. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  8. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost

Get the daily brief of stories like this at 6:30 every morning →