Qwen3.8-flash-next (iq2_xs)
strata model: qwen3.8-flash-next-iq2_xs engine: strata 0.1.33 context limit: 131,072 tokens gpu: rtx 4090, 24 gb system memory: 31 gb cpu: intel i9-14900k current settings - key/value cache: int8 - speculative decoding: enabled, 4-token window - expert cache: automatic - expert profile: enabled - e…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-02 19:21 · r/LocalLLM
Qwen3.8-flash-next (iq2_xs)