AWS just replaced round-robin with GPU-aware routing for LLM inference
AWS released a new Inference Gateway for Sagemaker Hyperpod. Instead of sending requests to the next available pod, it checks what’s actually happening on each one: queue depth, KV cache usage, prefix cache hits, loaded LoRA adapters, and current requests. That makes sense for LLM workloads. Two GP…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-10-11 15:44 · r/machinelearningnews
AWS just replaced round-robin with GPU-aware routing for LLM inference