GPU offload says Max, but only 54 of 65 layers loaded: how to check where your model actually runs
This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.
A 27B model at Q4 was generating 12-18 tokens per second on an RTX 3090. Same model, same settings a week earlier, it was doing 60-70. Nothing in the app looked broken. The model config still said GPU Offload: Max. That's the trap. Max doesn't mean all layers on the GPU. It means as many as the app…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-30 17:10 · DEV Community — AI
GPU offload says Max, but only 54 of 65 layers loaded: how to check where your model actually runs