AINewsnow

Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. ๐Ÿ˜ตโ€๐Ÿ’ซ Adjusting context will increase/decrease usage on both side. --tensor-split 499,501 = GPU1 12.5 GB, GPU2 15.โ€ฆ

Read the full story at r/LocalLLaMA โ†—

Timeline ยท 1 report

  1. 2026-09-27 20:18 ยท r/LocalLLaMA
    Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

More stories

  1. Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding โ€” MarkTechPost
  2. Faster prompt lookup drafting in llama.cpp โ€” Hacker News Front Page
  3. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine โ€” r/LocalLLM
  4. Updated from 3x3090(2x3090, 1x3090TI) to 2x5090 โ€” r/LocalLLaMA
  5. is switching from llama cpp to vllm worth it โ€” r/LocalLLaMA
  6. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy โ€” r/LocalLLaMA
  7. I ran the Laya model on llama.cpp to test the game of Snake, and the results were excellent. โ€” r/LocalLLM
  8. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second โ€” r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning โ†’