AINewsnow

Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. However pulled it out again when Qwen3.8 dropped, and it was much improved, partic…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-04 08:06 · r/LocalLLaMA
    Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4

More stories

  1. I turned an asymetric pair of Tesla V100s PCIe both (16 GB + 32 GB) into a surprisingly capable local LLM lab — 1.38k prompt tok/s, 40 decode tok/s with qwen3.8 27B Q6 and Q8... — r/LocalLLaMA
  2. Migrating from a R640 with passive cards — r/LocalLLaMA
  3. V100 16GB worth it?? — r/LocalLLM
  4. Sunsetting the NVIDIA Tesla P100 GPU on September 15, 2026 | What will happen to these P100, can we buy them? — r/LocalLLM
  5. When do you think we’ll get physical AGI? — r/singularity
  6. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  7. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  8. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →