AINewsnow

Ollama silently truncated my context window. A scanner for local LLM setups

This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.

On my MacBook Air, Ollama 0.34.4 served qwen3:1.7b with a 4,096-token window. The model supports 40,960. I sent a 7,000-token prompt and got HTTP 200 back, but the model had only seen the last 2,050 tokens. There was no error and no warning. That was the start of llm-doctor, a health check for loca…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-30 19:53 · DEV Community — AI
    Ollama silently truncated my context window. A scanner for local LLM setups

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  3. Introducing dots — OpenAI News
  4. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  5. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  6. OpenAI pauses AI model training after another agent bypasses network restrictions — InfoWorld AI
  7. Ollama now supports Jev-style decision models — Ollama Blog
  8. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →