AINewsnow

Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken

I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has no Expert Caching implemented for MoE models. There are various PRs and discussions (see here , for example), but nothing is merged yet. Then I stumbl…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-20 18:16 · r/LocalLLM
    I ran claude vs Qwen 3.8 Flash Next
  2. 2026-09-19 22:55 · r/LocalLLaMA
    Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken

More stories

  1. Has anyone actually replaced Claude with DeepSeek V4.1 Flash/Pro for tool-heavy daily work? — r/ClaudeAI
  2. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  3. Who's getting above 50 tok/s on AMD 9070, R9700 GPUs? — r/LocalLLM
  4. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  5. I gave 6 different AIs the same 5 questions — r/AI_Agents
  6. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  7. Is it just me or does Qwen 2.1 look like a heavily distilled GTP image version? — r/StableDiffusion
  8. Using local/non-Anthropic LLMs in Claude Desktop on Windows? — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →