Mellum2.1 MLX NVFP4 for 16 GB Macs
I converted JetBrains’ Mellum2.1 to MLX NVFP4: 6.84 GB, 64 tokens/s at 1K context on my M4. Would love feedback on real tasks Try it https://huggingface.co/imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-nvfp4 Benchmarks/code https://github.com/imaddde867/mellum-mlx
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-11 17:50 · r/LocalLLM
Mellum2.1 MLX NVFP4 for 16 GB Macs