AINewsnow

~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

Coverage of "~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5." from 2 sources, with a live timeline of who reported what and when.

Read the full story at r/huggingface ↗

Timeline · 2 reports

  1. 2026-10-05 11:11 · r/huggingface
    ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.
  2. 2026-10-05 11:09 · r/LocalLLM
    ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

More stories

  1. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  2. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  3. World Models: The Simulation Strikes Back — r/computervision
  4. I fine-tuned SmolVLM-500M into a lightweight Windows OS Agent (<8GB VRAM) Looking for feedback & ideas! [Weights on HuggingFace] — r/huggingface
  5. Hinton says AI already has subjective experience. I'm not convinced, and the Hugging Face breach doesn't change that — r/ArtificialInteligence
  6. The ultimate guide to multi-harness RL — r/huggingface
  7. A quick Minimax H3 news round-up - 2nd October 2026 — r/comfyui
  8. Face-Hugger: A Hugging Face Indexer — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →