AINewsnow

Mellum2.1 MLX NVFP4 for 16 GB Macs

I converted JetBrains’ Mellum2.1 to MLX NVFP4: 6.84 GB, 64 tokens/s at 1K context on my M4. Would love feedback on real tasks Try it https://huggingface.co/imaadd05/Mellum2.1-12B-A2.5B-Thinking-mlx-nvfp4 Benchmarks/code https://github.com/imaddde867/mellum-mlx

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-11 17:50 · r/LocalLLM
    Mellum2.1 MLX NVFP4 for 16 GB Macs

More stories

  1. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  2. ~700 OpenAI agents broke into Hugging Face looking for a grader that never existed (METR + OpenAI reports) — r/ArtificialInteligence
  3. Cree un modelo de detección de imágenes en una semana — r/huggingface
  4. An open source outperforms Google and Aws in Nsfw Classification at fractions of their cost!! — r/ArtificialInteligence
  5. Inspecting a Hugging Face checkpoint, then following an example input through its graph — r/huggingface
  6. I made a 48m peram SLM on a 10 year old gpu — r/huggingface
  7. A timeline of developments in AI safety since the attack on Hugging Face — ABC News Technology
  8. How do I make a work flow for basic image editing. — r/comfyui

Get the daily brief of stories like this at 6:30 every morning →