AINewsnow

Run Qwen3.8+Flash-Next and tiny models on Apple Silicon up to 3x faster

Maybe you'll like it? I hope I get to use my self-promotion credit a tiny little bit here after being in the community so long haha. I was the top of MLX.fast for a while and remain the winner on chips below M5. If you have capacity to contribute further enhancements I'd love that https://github.co…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-26 21:10 · r/LocalLLaMA
    Run Qwen3.8+Flash-Next and tiny models on Apple Silicon up to 3x faster

More stories

  1. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  2. Compressing Streaming Neural Audio Encoders via Latent-Space Distillation — Apple Machine Learning Research
  3. MacBook Pro M5 Max LSE LLM running an AMD Radeon AI PRO R9700 over Thunderbolt 5 in a Razer enclosure — r/LocalLLaMA
  4. Uncensored LLM models for local use — r/LocalLLM
  5. Can Apple Home’s AI camera features outsmart Amazon’s and Google’s? I put them to the test — The Verge AI
  6. Mark Zuckerberg predicts Muse will become a "personal superintelligence" for billions. — Axios AI+
  7. I got Qwen-Image 2.1 (f16) text-to-image and image editing running on an M1 Max ~24s at 512x512 — r/StableDiffusion
  8. Bad Apple made by Opus 5.5 — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →