AINewsnow

Inference time grows dramatically after a few runs.

This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.

Laptop with i712700H, RTX 4050 6GB, 16GB DDR5, 1TB NVMe SSD. I was testing bigger models that I thought my machine couldn't handle but the newer ComfyUI versions with dynamic VRAM made it not only possible but also quite fast. The models in question are Qwen Image and Flux 2 Klein 9B, FP8 versions…

Read the full story at r/comfyui ↗

Timeline · 1 report

  1. 2026-08-20 11:58 · r/comfyui
    Inference time grows dramatically after a few runs.

More stories

  1. Best model/workflow for face and body consistency? — r/comfyui
  2. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  3. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLM
  4. Is Qwen 3.8 Flash Next usable on M1 Ultra 64GB ? — r/LocalLLM
  5. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  6. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  7. Qwen Developers on X: "Qwen-Image 2.1 is going open source" — r/StableDiffusion
  8. Ternary Bonsai 2 27B — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →