AINewsnow

Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me

This story is from 2026-10-10. It is preserved in the archive; the latest stories are on the live feed.

My local RAG bot answered 46% of my test questions correctly. Llama 3.1 8B on Ollama, a 128K context window on the model card, retrieved chunks that I had checked by hand. The right paragraph was in the prompt every single time. The model just never saw it. Ollama's num_ctx was set to 2048 tokens,…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-10 16:52 · DEV Community — AI
    Ollama num_ctx Truncated 287 of 400 Prompts and Never Told Me

More stories

  1. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  2. Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects — r/LocalLLaMA
  3. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  4. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  5. Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM — r/LocalLLM
  6. Created laya : Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support — r/LocalLLaMA
  7. Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec? — r/LocalLLaMA
  8. Qwen3.8-27B on a single 3090: 140 tok/s on code with a custom megakernel — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →