AINewsnow

Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max

I tested Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max this is the test as article written by ChatGPT ,I just run the commands, I'm sharing the results here in case someone is interested. TL;DR Tested locally on a MacBook Pro M2 Max, 96GB unified memory , using MLX: Model Speed Peak…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-29 11:18 · r/LocalLLaMA
    Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide
  2. 2026-09-28 16:25 · r/LocalLLM
    Qwen 3.8 27B Q4/Q6/Q8 vs Qwen 3.8 Flash-Next on a 96GB M2 Max

More stories

  1. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  2. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s — r/LocalLLM
  5. If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option? — r/LocalLLaMA
  6. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  7. Qwen 3.8 is a workhorse — r/LocalLLaMA
  8. OpenAI pauses AI model training after another agent bypasses network restrictions — InfoWorld AI

Get the daily brief of stories like this at 6:30 every morning →