AINewsnow

PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig

I run a rack of second-hand Pascal and Volta cards on PCIe x4 risers, and I've been building an engine for exactly that kind of hardware: its own quant format (PXQ), its own kernels for P100 / GTX 1080 Ti / V100, and a launcher that only asks which cards and which model. This is the first release t…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-21 21:18 · r/LocalLLM
    PXA v2026.09.20 — my inference engine for old Teslas (P100 / V100 / 1080 Ti): Gemma 4 MoE, tensor split on by default, and ahead of stock llama.cpp on every cell on my rig

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  3. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  4. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  5. Who's getting above 50 tok/s on AMD 9070, R9700 GPUs? — r/LocalLLM
  6. Ternary Bonsai 2 (27B) fails to load in LM Studio and oMLX. I made fixes for both (GGUF PQ2_0/PTQ1_0 + MLX 2-bit) — r/LocalLLM
  7. You can use any LLM just like JEV — r/LocalLLaMA
  8. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →