Advice on model routing
Hi folks, I am interested in using exo and Ollama to route requests between 2 separate pods. The primary one is 192GB doing 125bq4 256k context. The other is 64GB doing 35bq4 128k context. I want the clients to autoroute without selecting models. Back end is a mix of NVIDIA and Rocky Linux for the…
Read the full story at r/MLQuestions ↗
Timeline · 1 report
- 2026-09-24 12:06 · r/MLQuestions
Advice on model routing