AINewsnow

Qwen3.8 27B runs at ~480 tok/s on a single 5090

Been doing some testing, swapped NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) for QUASAR's QAT full-NVFP4 checkpoint, with DFlash2 embedded. Running on the same hardware and engine flags: metric official this JSON output 400 tok/s 484 tok/s prose 176 tok/s 204 tok/s prefill 11.4k tok/…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-01 16:30 · r/LocalLLM
    Qwen3.8 27B runs at ~480 tok/s on a single 5090

More stories

  1. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns — New York Times Technology
  4. Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
  5. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  6. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  7. Introducing dots — OpenAI News
  8. Ollama now supports Jev-style decision models — Ollama Blog

Get the daily brief of stories like this at 6:30 every morning →