Ornith seems to be better.
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: On an M3 Ultra, Ornith-1.5-35B-A3B (4-bit MLX) decodes 4.6× faster than Qwen3.8-27B (8-bit MLX) and scores slightly higher on a small hard eval. It also beats Qwen3.8-27B with speculative decoding, while running autoregressive. Setup Mac Studio, M3 Ultra, 256 GB unified memory mlx-lm 0.31.3…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 10:47 · r/LocalLLM
Ornith seems to be better.