AINewsnow

Qwen3.6 35B-A3B optimized for a 4 GB GTX 1050 Ti Mobile: up to 2× prefill and 3× decode vs llama.cpp

I optimized Strata’s Qwen3.6 35B-A3B inference path for a Pascal GTX 1050 Ti Mobile with 4 GB VRAM. The main changes include native quantized expert execution, packed-IQ DP4A MMQ, layer-major prefill, phase-specific VRAM reuse, concurrent CPU/GPU MoE execution, adaptive expert caching, and MTP opti…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-09 16:50 · r/LocalLLM
    Qwen3.6 35B-A3B optimized for a 4 GB GTX 1050 Ti Mobile: up to 2× prefill and 3× decode vs llama.cpp

More stories

  1. feat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects — r/LocalLLaMA
  3. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  4. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  5. Suggestion for llama.cpp configuration — r/LocalLLM
  6. How to Fine-Tune Llama 3 for Custom Tool Calling with Unsloth in Python — Machine Learning Mastery
  7. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  8. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →