Make Volta Fast Again
For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples, but I decided to sh…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-25 17:02 · r/LocalLLaMA
Make Volta Fast Again