AINewsnow

PXA v3.1: Qwen3.8 Flash-Next at 31 t/s on a single Tesla P100, 55 t/s on two V100s (open source, beats Strata on prose + prefill)

https://preview.redd.it/smmxfg9d4ruh1.jpg?width=1800&format=pjpg&auto=webp&s=28b2129c598488c90cad948d6bf0fe0b472390ba https://preview.redd.it/4qnolw5c4ruh1.png?width=1695&format=png&auto=webp&s=cf5872bafa4b40339d13ab2817ab70b2b1779ed2 I'm the dev of PXA, an open-source inference engine built for ol…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-11 02:33 · r/LocalLLM
    PXA v3.1: Qwen3.8 Flash-Next at 31 t/s on a single Tesla P100, 55 t/s on two V100s (open source, beats Strata on prose + prefill)

More stories

  1. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  2. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude — The Guardian AI
  3. Philadelphia police receive false homicide tip from Anthropic AI model — The Hill Technology
  4. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  5. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog
  6. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  7. Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE AI
  8. Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash — Cloudflare Blog — AI

Get the daily brief of stories like this at 6:30 every morning →