AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
Faster prompt processing on CPU.
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-31 18:53 · r/LocalLLaMA
AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp