AINewsnow

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

a while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~ 430 tok/s prompt processing, and…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-27 13:17 · r/LocalLLM
    Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

More stories

  1. Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding — MarkTechPost
  2. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  4. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  5. Advice on models for RAG use case — r/LocalLLaMA
  6. Targeting Applied AI / ML Engineer roles. Need ruthless feedback on my architecture and metrics. — r/deeplearning
  7. 42x Faster Prompt Lookup Drafting in llama.cpp — r/LocalLLaMA
  8. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →