AINewsnow

M2 Mac ultra128gb Qwen flash next

I am trying to get good speeds for Qwen flash next but I also need some ram for the system. Right now gpu is at 96gb around more I cannot use. Before I was getting 50 token/s from omlx and 60 tokens/s from mlx serve but lower context like 32k. I had ngram on ssd now where my system has more ram for…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-20 18:16 · r/LocalLLM
    I ran claude vs Qwen 3.8 Flash Next
  2. 2026-09-19 22:55 · r/LocalLLaMA
    Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken
  3. 2026-09-19 01:20 · r/LocalLLM
    M2 Mac ultra128gb Qwen flash next

More stories

  1. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  4. I gave 6 different AIs the same 5 questions — r/AI_Agents
  5. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Is it just me or does Qwen 2.1 look like a heavily distilled GTP image version? — r/StableDiffusion
  7. Using local/non-Anthropic LLMs in Claude Desktop on Windows? — r/ClaudeAI
  8. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →