Built a 3-model local agent swarm on a single 16GB card, partially offloaded to 32GB RAM — what would YOU do with 2x Ling tiny + Qwen3.8 oversight?
Hardware: RTX 5060 Ti 16GB, 32GB RAM, Ryzen 7 8700F (WSL2 Ubuntu). Everything local. The swarm (all three resident concurrently, ~14.5/16.3GB VRAM): - Overseer — Qwen3.8-27B "Mirai S" (alesha-pro 2.4-bit fork): 128k ctx, q4_0 KV, text-only, no MTP. ~12GB VRAM. (Alternate lead staged but untested: T…
Read the full story at r/LocalLLM ↗