PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig
I run a rack of second-hand Pascal and Volta cards on PCIe x4 risers, and I've been building an engine for exactly that kind of hardware: its own quant format (PXQ), its own kernels for P100 / GTX 1080 Ti / V100, and a launcher that only asks which cards and which model. This is the first release t…
Read the full story at r/LocalLLM ↗