htop for LLM inference just went multi-GPU π
Your LLM is using 14 GB of VRAM. 14 GB of what? π Weights? KV cache? CUDA overhead? One GPU or two? With tensor parallelism it gets even messier. vLLM shows you: EngineCore Worker_TP0 Worker_TP1 But that's not three workloads. It's ONE model running across multiple GPUs. That's why LLM Inspector vβ¦
Read the full story at r/learnmachinelearning β
Timeline Β· 2 reports
- 2026-09-30 09:14 Β· r/learnmachinelearning
htop for LLM inference just went multi-GPU π - 2026-09-30 09:13 Β· r/learnmachinelearning
htop for LLM inference just went multi-GPU π