Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. However pulled it out again when Qwen3.8 dropped, and it was much improved, partic…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-04 08:06 · r/LocalLLaMA
Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4