Ran qwen3.8 flash next coder in strata fork on 16gb ram and vram
I built a strata fork for using qwen3 8 flash next coder with 16gb ram and not 32gb using nvme and I have around 10 to 20 tokens per second let me know if I should put it on GitHub. I'll post a picture when I'm able to.
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-07 04:58 · r/LocalLLM
Ran qwen3.8 flash next coder in strata fork on 16gb ram and vram