AINewsnow

Advice on model routing

Hi folks, I am interested in using exo and Ollama to route requests between 2 separate pods. The primary one is 192GB doing 125bq4 256k context. The other is 64GB doing 35bq4 128k context. I want the clients to autoroute without selecting models. Back end is a mix of NVIDIA and Rocky Linux for the…

Read the full story at r/MLQuestions ↗

Timeline · 1 report

  1. 2026-09-24 12:06 · r/MLQuestions
    Advice on model routing

More stories

  1. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  2. NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development — NVIDIA Blog
  3. How far behind Nvidia is Huawei? — Epoch AI
  4. NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time — MarkTechPost
  5. notjev – Turn any local LLM into a Jev-style decision engine (one token + logprobs) — r/LocalLLM
  6. How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows — Hugging Face Blog
  7. At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia — NVIDIA Blog
  8. Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing — NVIDIA Technical Blog

Get the daily brief of stories like this at 6:30 every morning →