How to have more context without loosing speed?
This story is from 2026-08-19. It is preserved in the archive; the latest stories are on the live feed.
I am running Qwen3.8 27b q2 with 12 gb vram and in the desktop app it says that I can only have context 4096 or it will use my RAM, and when it does that it is super slow. Is there a way to have the same speed even with larger context? Please I need a magical fix π
Read the full story at r/LocalLLM β
Timeline Β· 1 report
- 2026-08-19 17:04 Β· r/LocalLLM
How to have more context without loosing speed?