AINewsnow

Deepseek V4 Flash 0731 on m5 max 128gb

I have an m5 max mac studio 128gb on order. For my specific workload and with a few openrouter credits I have found that deepseek v4 flash 0731 works the best for me, even better than qwen 3.8 27b (although that was pretty good) (couldn’t run qwen 3.8 flash for excessive content moderation reasons…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-28 11:08 · r/LocalLLM
    Deepseek V4 Flash 0731 on m5 max 128gb

More stories

  1. Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much? — r/LocalLLaMA
  2. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  3. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/artificial
  4. Qwen 3.8 is a workhorse — r/LocalLLaMA
  5. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents
  6. When is the next generation of "B tier" models releasing? — r/LocalLLaMA
  7. One key for claude, gpt, gemini, and deepseek in my coding tools — r/ChatGPTCoding
  8. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →