AINewsnow

I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

I've been investigating hardware-accelerated llama.cpp inference on a Galaxy S24 Ultra (Snapdragon 8 Gen 3) from ordinary F-Droid Termux without root. I initially set out to get the Adreno 750 OpenCL backend working. That now works reproducibly and executes real GPU kernels, although generic OpenCL…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-04 03:21 · r/LocalLLM
    I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

More stories

  1. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant — r/LocalLLaMA
  5. FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
  6. Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality — r/LocalLLM
  7. The Rise of Overfit Inference Engines — r/LocalLLaMA
  8. Flash next rig born from mining parts. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →