AINewsnow

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.

Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini I went with the recommended settings for the best generation quality and ran a few benchmarks: temperature: 0.7 top_p: 0.95 top_k: 40 reasoning_effort: high I must say, the outcome is no…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-12 19:00 · r/LocalLLaMA
    Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

More stories

  1. A closer look at the upcoming Siri AI-powered home hub, a key pillar of Apple's strategy for the home; sources: Apple started cutting Fitness+ staff (Mark Gurman/Bloomberg) — Techmeme
  2. Gemini Joins the Hacker Club — Wall Street Journal Technology
  3. Apple’s Home AI Hub Details; Apple Fitness+ Layoffs and iPhone Duo Apple Pencil — Bloomberg AI
  4. ComfyUI on Apple Silicon: no MLX, no fp8, 600-second kernel builds. So I built my own launcher — a personal project I'm sharing in case it helps someone. — r/comfyui
  5. [Release] Nirvana Code: A single-binary Rust LLM engine built from the metal up for Apple Silicon (Metal 3, Persistent Prefix Cache, Speculative Decoding, Dual GGUF + MLX) — r/LocalLLM
  6. He’s the Face of AI Doomsday Fears — Wall Street Journal Technology
  7. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  8. Vibecoded a Research paper app — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →