VLLM on Mac Nvidia 5090 GPU
So after getting llamacpp working at native speeds on my 5090 I'm working on the next project concurrently. VLLM and PyTorch with native cuda on Mac. This is requiring a different path and so far the single context is much slower but the reason I started was for batching and in that test it's worki…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-22 14:20 · r/LocalLLM
VLLM on Mac Nvidia 5090 GPU