Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me
This story is from 2026-10-10. It is preserved in the archive; the latest stories are on the live feed.
My local RAG bot answered 46% of my test questions correctly. Llama 3.1 8B on Ollama, a 128K context window on the model card, retrieved chunks that I had checked by hand. The right paragraph was in the prompt every single time. The model just never saw it. Ollama's num_ctx was set to 2048 tokens,…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-10 16:52 · DEV Community — AI
Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me