Best current R9700 inference engine?
There are way too many forks to keep track of, so I've gotten lost. As far as I can tell, Radiance VLLM is best for models that fit in GPUs while some form of llama.cpp is probably best for MOE RAM-spill? My specific current goal is to run GLM5.3-Flash over 6 R9700s + system RAM but also looking to…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-08 10:16 · r/LocalLLaMA
Best current R9700 inference engine?