Tesla V100 32GB + Qwen3.8-27B at 23.6 tok/s and 256K context — cooled by a blower mounted with Velcro
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
I’ve been building a small heterogeneous local-AI lab, and “Team Green” has turned into the strangest useful machine in it. The system: Ryzen 9 9950X 48 GB DDR5 HPE/NVIDIA Tesla V100 PCIe 32 GB HBM2 ECC Ubuntu 24.04 llama.cpp build 10499 Qwen3.8-27B Q3_K_M, 12.86 GiB All 66/66 layers offloaded to t…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-24 12:36 · r/LocalLLM
Tesla V100 32GB + Qwen3.8-27B at 23.6 tok/s and 256K context — cooled by a blower mounted with Velcro