Long prompts on an M1 Max: Splash-M1 vs MTPLX vs oMLX vs TensorFold, 3-turn chats from 2K to 256K tokens (Qwen3.8-27B)
TL;DR (M1 Max 64GB, Qwen3.8-27B, long input with short answers): Splash-M1 decodes fastest up to 64K and uses the least memory. My MTPLX M1 fork reads long prompts fastest, so its whole 3-turn conversation was the shortest at every length (15% shorter than Splash at 128K), at the cost of much highe…
Read the full story at r/LocalLLM ↗