qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp
Qwen Flash Next now uses less VRAM
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-03 06:18 · r/LocalLLaMA
qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp