AINewsnow

PSA for --n-cpu-moe users on NVIDIA: check your memory clock during decode. Mine was sitting at 810 MHz. Locking clocks gave +40% on one GPU and 3x on two.

This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.

TL;DR During MoE offload decode the GPU waits on the CPU most of each token, so utilization reads 20 to 40 percent. The NVIDIA driver reads that as idle and drops the card to P5: about 480 MHz core and 810 MHz memory, down from 7601. Decode is memory-bound, so it falls with it. Prompt processing ke…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-08-21 01:36 · r/LocalLLM
    PSA for --n-cpu-moe users on NVIDIA: check your memory clock during decode. Mine was sitting at 810 MHz. Locking clocks gave +40% on one GPU and 3x on two.

More stories

  1. Amid growing AI fears, King Charles meets with industry leaders in Scotland — NPR Technology
  2. King Charles to press Nvidia, OpenAI, Anthropic leaders on AI safety at summit — CNBC Technology
  3. Huawei details AI accelerator roadmap, pulls in next-generation Ascend NPUs by several quarters — FP4 performance of the Ascend 960PR doubles expectations — Tom's Hardware
  4. Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors. — r/LocalLLaMA
  5. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  6. Huawei unveils latest tech to boost AI power in push to break China’s Nvidia reliance — South China Morning Post Tech
  7. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  8. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →