How would you go about using multiple models together from a singe router (?) or a end point ?
I have 3 machines, My main one can run Qwen3.8 Flash Next at 13-15 tps, i also have an MacMini 16GB whcih can run Orninth 9B or Gemma4 12B easily and i have a Pi5 8B that can run a 3B model well. I want to run an EndPoint/Router that is connected to the harness, that breaks down the task and distri…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-25 18:51 · r/LocalLLaMA
How would you go about using multiple models together from a singe router (?) or a end point ?