AINewsnow

Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?

I'm currently choosing between IQ3_XXS Qwen-3.8-Flash and Q8_0 27B. This month I don't have anything complex enough to justify either's potential so I've got a fairly bad read on how these two stack up in terms of intelligence. Have any of you compared the two enough to speak to which you've had a…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-02 17:22 · r/LocalLLaMA
    Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?

More stories

  1. Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
  2. Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
  3. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  4. I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. — r/LocalLLaMA
  5. Some topics are off limits in Chinese AI, researchers find — CBS News Technology
  6. What is your experience with bonsai 2 27b? — r/ArtificialInteligence
  7. Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print. — r/LocalLLaMA
  8. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →