AINewsnow

Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 GB M5 Max. The post Pe…

Read the full story at MarkTechPost ↗

Timeline · 2 reports

  1. 2026-09-03 07:08 · r/machinelearningnews
    Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
  2. 2026-09-03 06:57 · MarkTechPost
    Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

More stories

  1. A closer look at the upcoming Siri AI-powered home hub, a key pillar of Apple's strategy for the home; sources: Apple started cutting Fitness+ staff (Mark Gurman/Bloomberg) — Techmeme
  2. Gemini Joins the Hacker Club — Wall Street Journal Technology
  3. Apple’s Home AI Hub Details; Apple Fitness+ Layoffs and iPhone Duo Apple Pencil — Bloomberg AI
  4. ComfyUI on Apple Silicon: no MLX, no fp8, 600-second kernel builds. So I built my own launcher — a personal project I'm sharing in case it helps someone. — r/comfyui
  5. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  6. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  7. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  8. AI agents/automation suggestions for a solo biz — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →