Perplexity's Lily Beats MLX-LM by 1.35x Running Qwen3.6 on Apple Silicon
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
Perplexity open-sources Lily, a Metal-based inference engine tuned for Qwen3.6-35B-A3B that beats MLX-LM by 1.23x prefill and 1.35x decode on M5 Max.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-09-02 20:04 · AlphaSignal
Perplexity's Lily Beats MLX-LM by 1.35x Running Qwen3.6 on Apple Silicon