Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: I'm planning to spend around $100 on cloud GPUs to benchmark Qwen3.8-27B with a focus on questions that actually matter when running it locally: different quant levels/providers, 8-bit vs 16-bit KV cache, GGUF vs EXL3, context length tradeoffs, and token efficiency on coding/agentic workload…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-24 19:00 · r/LocalLLaMA
Planning to spend ~$100 benchmarking differnet Qwen3.8-27B quants and kv cache and looking for input before I start