FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
arXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request ro…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.AI
FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving