AINewsnow

I have single 3090 and 32GB of ram are there any engines like strata, ninfer that run Qwen/Qwen3.6-35B-A3B at Q8 fast on my hardware

strata with qwen flash Q2 is pretty retarded and 27b at q4 is also retarded so I am wondering if anyone knows a engine that runs the model Qwen/Qwen3.6-35B-A3B fast on my hardware?

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-07 21:05 · r/LocalLLaMA
    I have single 3090 and 32GB of ram are there any engines like strata, ninfer that run Qwen/Qwen3.6-35B-A3B at Q8 fast on my hardware

More stories

  1. Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M) — r/LocalLLaMA
  2. Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? — r/LocalLLaMA
  3. China’s open-weight AI models are winning global users. Who is capturing the value? — South China Morning Post Tech
  4. Hunyuan image 3 vs qwen 2.1 — r/StableDiffusion
  5. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  7. A Very Strange GPU Stall Situation — r/comfyui
  8. Stepping away from Benchmarks and Code, what models are you using for Creating writing projects — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →