AINewsnow

Sub-15ms Local Decisions: Running Laya Models with Native MLX on Apple Silicon

This story is from 2026-10-10. It is preserved in the archive; the latest stories are on the live feed.

TL;DR Stop burning tokens and introducing multi-second cloud latency just to make structured routing decisions. laya-mlx is a native Apple MLX runtime for Laya typed decision models that clocks in at an astonishing 7–14 ms on an M3 Max. By stripping away autoregressive text generation and heavy PyT…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-10 01:52 · DEV Community — AI
    Sub-15ms Local Decisions: Running Laya Models with Native MLX on Apple Silicon

More stories

  1. I built an open source framework for building RL environments. Named "Seahaven" after the fake town in The Truman Show. — r/LocalLLaMA
  2. Nara Baby introduced paid plans, so I tried building our own baby tracker with Claude Code — r/ClaudeAI
  3. I run a persistent local agent on a 16GB Air. Qwen3-8B on the GPU, a second brain on the Neural Engine, no cloud. — r/LocalLLM
  4. Claude created my dream game, and got approved for Apple iOS store! — r/ClaudeAI
  5. Whallm 1.1.11: Swift1.5-Qwen3.8 support, and Qwen3.8 now runs at 15–17 tok/s decode and 500–600 tok/s prefill at 16K on a 64 GB Mac — r/LocalLLM
  6. Switch Billing - Lose Resets? — r/OpenAI
  7. What would make you actually use a personal AI assistant everyday? — r/artificial
  8. My ComfyUI Nodes and Workflows - Krea 2 (Turbo and Raw), Z-Image (Turbo and Base), MiniMax Music 3, Image2Text and LLM Chat (with Tools), Torch, Apple MLX and Cloud — r/comfyui

Get the daily brief of stories like this at 6:30 every morning →