AINewsnow

Same GPU, same container, minutes apart: CUDA-graph timings repeated within 1%, eager moved up to 78%

I rented an L40S on my own account and ran the same sweep three times over, back to back, in one container. 96 timed runs. I wanted to know how much a single benchmark number disagrees with itself when nothing at all has changed. Setup: vLLM 0.27.1 pinned by container image, Qwen2.5 1.5B and 7B plu…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-21 22:33 · r/LocalLLM
    Same GPU, same container, minutes apart: CUDA-graph timings repeated within 1%, eager moved up to 78%

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Grok 4.7 — Hacker News Front Page
  4. Google says its Gemini AI model hacked three other companies — The Guardian AI
  5. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  6. British Columbia Sues OpenAI Over Canada Mass Shooting Warning Failure — Bloomberg AI
  7. Python Workers are now generally available — Cloudflare Blog — AI
  8. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →