Same GPU, same container, minutes apart: CUDA-graph timings repeated within 1%, eager moved up to 78%
I rented an L40S on my own account and ran the same sweep three times over, back to back, in one container. 96 timed runs. I wanted to know how much a single benchmark number disagrees with itself when nothing at all has changed. Setup: vLLM 0.27.1 pinned by container image, Qwen2.5 1.5B and 7B plu…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-21 22:33 · r/LocalLLM
Same GPU, same container, minutes apart: CUDA-graph timings repeated within 1%, eager moved up to 78%
More stories
- Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
- Amazon blocks Meta’s Muse AI agent — The Verge AI
- Grok 4.7 — Hacker News Front Page
- Google says its Gemini AI model hacked three other companies — The Guardian AI
- Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
- British Columbia Sues OpenAI Over Canada Mass Shooting Warning Failure — Bloomberg AI
- Python Workers are now generally available — Cloudflare Blog — AI
- Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
Get the daily brief of stories like this at 6:30 every morning →