AINewsnow

Best models for Mac Studio M2 Ultra 128GB?

As of October 2026, what are the best models for agentic coding (using OpenCode as my harness) and agentic chat (using OpenWeb UI as my harness) that can hit at least ~40 tok/sec? Can be separate and would like native CoT support. Messing around with Qwen 3.6 and 3.8 Flash but wondering if there ar…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-04 05:14 · r/LocalLLM
    Best models for Mac Studio M2 Ultra 128GB?

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  5. I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA
  6. Use Qwen-Image-2.1-viggle-turbo to generate character sheets in ComfyUI — r/StableDiffusion
  7. Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it — r/machinelearningnews
  8. I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →