How do I get my main local LLM use a secondary local LLM for simpler tasks?
My main setup is a M5 Max with 64GB. I usually run Qwen3.8-27B in oMLX and talk to it through the Hermes tui. I pulled out my older Dell, it has a RTX 3080 Laptop GPU with 16GB of VRAM. I put LM Studio on that and tried out Qwen3.8-9B, it runs on there well. That made me wonder if there's a way to…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-11 03:00 · r/LocalLLM
How do I get my main local LLM use a secondary local LLM for simpler tasks?