I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)
I've been investigating hardware-accelerated llama.cpp inference on a Galaxy S24 Ultra (Snapdragon 8 Gen 3) from ordinary F-Droid Termux without root. I initially set out to get the Adreno 750 OpenCL backend working. That now works reproducibly and executes real GPU kernels, although generic OpenCL…
Read the full story at r/LocalLLM ↗