Run Qwen 3.8 27b on the Apple Neural Engine at 7 watts on a Mac
Posted here about an inference engine I was working on to get Qwen 3.8 27b working with the Apple Neural engine at full context. I updated to OSX 27 and had to rework the engine, but I've now put it up publicly for people to use as well as the models on Hugging Face. There is also a hybrid GPU + AN…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-25 10:03 · r/LocalLLM
Run Qwen 3.8 27b on the Apple Neural Engine at 7 watts on a Mac