Apple Neural Engine Reverse-Engineered for Qwen 3.8 27B Inference
This story is from 2026-08-21. It is preserved in the archive; the latest stories are on the live feed.
A developer reverse-engineered Apple's Neural Engine to run Qwen 3.8 27B FP16 at 7-8 tok/s using 7W, while Mac users report 17-20 tok/s via LM Studio and RTX Pro 4000 SFF Blackwell users seek speed benchmarks.
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-08-22 12:16 · r/LocalLLM
Any of you running Qwen 3.8 27B on an RTX Pro 4000 SFF Blackwell? - 2026-08-22 04:48 · r/LocalLLM
people running Qwen 3.8 27B on apple silicon… whats your best token generation speed and how did you attain it? - 2026-08-21 12:02 · r/LocalLLM
Running Qwen 3.8 27b FP16 on the Apple Neural Engine - 7 Watts of power to run a FP16 model @ 7 tok/s