AINewsnow

Double GPU setup, How to get even compute distribution?

I got a 9070xt and a 9070 double GPU setup. I am pretty new to the local LLM usage and I am trying it out with unsloth utilizing qwen 3.8 27B IQ4_XS. I am using "-tensor-split" and "--split-mode layer" flags to control how the model is distributed among the GPUs. For the VRAM usage, these flags wor…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-04 10:07 · r/LocalLLM
    Double GPU setup, How to get even compute distribution?

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  5. ComfyUI Qwen image 2.1 Enhancer (Two nodes) — r/StableDiffusion
  6. I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA
  7. Use Qwen-Image-2.1-viggle-turbo to generate character sheets in ComfyUI — r/StableDiffusion
  8. Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it — r/machinelearningnews

Get the daily brief of stories like this at 6:30 every morning →