AINewsnow

Long prompts on an M1 Max: Splash-M1 vs MTPLX vs oMLX vs TensorFold, 3-turn chats from 2K to 256K tokens (Qwen3.8-27B)

TL;DR (M1 Max 64GB, Qwen3.8-27B, long input with short answers): Splash-M1 decodes fastest up to 64K and uses the least memory. My MTPLX M1 fork reads long prompts fastest, so its whole 3-turn conversation was the shortest at every length (15% shorter than Splash at 128K), at the cost of much highe…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 12:23 · r/LocalLLM
    Long prompts on an M1 Max: Splash-M1 vs MTPLX vs oMLX vs TensorFold, 3-turn chats from 2K to 256K tokens (Qwen3.8-27B)

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  3. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  4. OpenAI pauses AI training, launches ‘extensive’ review after multiple rogue agent incidents — Mint AI
  5. Anthropic warns of ‘existential risks to humanity’ in IPO prospectus — Financial Times AI
  6. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  7. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  8. OpenAI launches Dots, its Muse competitor — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →