Ollama silently truncated my context window. A scanner for local LLM setups
This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.
On my MacBook Air, Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window. The model supports 40,960. I sent a 7,000-token prompt and got HTTP 200 back, but the model had only seen the last 2,050 tokens. There was no error and no warning. That was the start of llm-doctor, a health check for loca…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-30 19:53 · DEV Community — AI
Ollama silently truncated my context window. A scanner for local LLM setups