Half a Million Tokens of Context Still Needs an Index
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
A thread on r/LocalLLaMA ran Qwen 3.8 27B with a 524,000-token context window on two RTX 3090s this week, holding 60 to 88 tokens per second. Two consumer GPUs, half a million tokens of context, one box. Five years ago that sentence would have been science fiction. Every milestone like this revives…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-10 02:41 · DEV Community — AI
Half a Million Tokens of Context Still Needs an Index