AINewsnow

Old datacenter GPUs aren't dead: PXA v3 runs 27B at 97 t/s on two V100s, 86 t/s on ONE, Gemma 4 at 169 t/s, and a 124B-class MoE on a single P100 + RAM (free, open engine + one-click GUI)

Hey r/LocalLLaMA 👋 We're PXA Network . We build PXA , an inference engine tuned for Pascal and Volta GPUs (P100, V100, P40, 1080 Ti), the $100–300 eBay cards everyone calls too old. PXA v3 just shipped, and everything runs through PXA Control , our one-click web app. ⚡ Speed (PXA v3, measured on o…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-06 02:54 · r/LocalLLM
    Old datacenter GPUs aren't dead: PXA v3 runs 27B at 97 t/s on two V100s, 86 t/s on ONE, Gemma 4 at 169 t/s, and a 124B-class MoE on a single P100 + RAM (free, open engine + one-click GUI)

More stories

  1. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  2. Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Sam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AI — Politico Technology
  4. OpenAI safety employee resigns, claiming the company’s ‘culture is broken’ — TechCrunch AI
  5. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  6. Supercharge regulated workloads with Claude Code and Amazon Bedrock — AWS Machine Learning Blog
  7. can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
  8. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →