Qwen3.6 35B-A3B optimized for a 4 GB GTX 1050 Ti Mobile: up to 2× prefill and 3× decode vs llama.cpp
I optimized Strata’s Qwen3.6 35B-A3B inference path for a Pascal GTX 1050 Ti Mobile with 4 GB VRAM. The main changes include native quantized expert execution, packed-IQ DP4A MMQ, layer-major prefill, phase-specific VRAM reuse, concurrent CPU/GPU MoE execution, adaptive expert caching, and MTP opti…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 16:50 · r/LocalLLM
Qwen3.6 35B-A3B optimized for a 4 GB GTX 1050 Ti Mobile: up to 2× prefill and 3× decode vs llama.cpp