Ok how to actually learn vLLM ?
Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / mor…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 08:18 · r/LocalLLaMA
Ok how to actually learn vLLM ?