dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode
Note not my work, but something i found and wanted to share so hopefully more people can push this along even further. https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt I basically run 2 x 7900 xtx on a consumer pc. This repo takes qwen 3.8 Q8 and optimizes it to run on this setup. My decode…
Read the full story at r/LocalLLaMA ↗