AINewsnow

GLM-5.3-Flash abliterated MLX 4-bit on mlx-serve, M3 Ultra 256 GB

Machine: Mac Studio M3 Ultra, 256 GB. Weights: grant-ai/GLM-5.3-Flash-Abliterated-MLX-4bit. Engine: mlx-serve 26.10.1 (MLX 0.32.3). Context 131072. The served copy is a view of that repo: fused conv1d split into q/k/v, forget-gate names lifted, vision tensors dropped. The source files are unchanged…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-11 11:47 · r/LocalLLM
    GLM-5.3-Flash abliterated MLX 4-bit on mlx-serve, M3 Ultra 256 GB

More stories

  1. My test of GLM 5.3 Flash Q2 on DGX Spark — r/LocalLLM
  2. Same GLM-5.3-Flash model, four provider endpoints: results from our physics arena — r/LocalLLM
  3. Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp — r/LocalLLM
  4. Mistral’s new Large 4 trails some Chinese open models in independent tests — Tom's Hardware
  5. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  6. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  7. Microsoft CEO Nadella Calls for ‘Emergency Brake’ on Advanced AI — Bloomberg AI
  8. Microsoft unveils Microsoft-Decision-1, a fast decision-scoring model trained on Qwen3.5-9B, and says it will soon rebase it on MAI, OpenAI, and other models (Achint Srivastava/Command Line) — Techmeme

Get the daily brief of stories like this at 6:30 every morning →