AINewsnow

Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal. What i run: - model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bp…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-02 11:19 · r/LocalLLaMA
    Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

More stories

  1. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  4. Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
  5. OpenAI announces ‘dots’ agent after scrapping launch of new AI model over safety concerns — The Guardian AI
  6. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  7. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  8. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog

Get the daily brief of stories like this at 6:30 every morning →