AINewsnow

~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

Stack Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4 Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding W4A8 prefill on Blackwell Main engine flags: ./build/stra…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-05 11:09 · r/LocalLLM
    ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

More stories

  1. Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
  2. World Models: The Simulation Strikes Back — r/computervision
  3. I fine-tuned SmolVLM-500M into a lightweight Windows OS Agent (<8GB VRAM) Looking for feedback & ideas! [Weights on HuggingFace] — r/huggingface
  4. Hinton says AI already has subjective experience. I'm not convinced, and the Hugging Face breach doesn't change that — r/ArtificialInteligence
  5. The ultimate guide to multi-harness RL — r/huggingface
  6. A quick Minimax H3 news round-up - 2nd October 2026 — r/comfyui
  7. Face-Hugger: A Hugging Face Indexer — r/huggingface
  8. Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →