AINewsnow

Whallm 1.1.11: Swift1.5-Qwen3.8 support, and Qwen3.8 now runs at 15–17 tok/s decode and 500–600 tok/s prefill at 16K on a 64 GB Mac

Whallm is a free, open-source macOS app that runs large MoE models on Apple Silicon by streaming experts from your SSD. New in 1.1.11 Swift1.5-Qwen3.8-Flash-Next support , with text and image input. Decode: ~9 → 15–17 tok/s. MTP is now on by default. Prefill: ~146 → 500–600 tok/s at 16K input. Peak…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-09 15:27 · r/LocalLLM
    Whallm 1.1.11: Swift1.5-Qwen3.8 support, and Qwen3.8 now runs at 15–17 tok/s decode and 500–600 tok/s prefill at 16K on a 64 GB Mac

More stories

  1. I built an open source framework for building RL environments. Named "Seahaven" after the fake town in The Truman Show. — r/LocalLLaMA
  2. Switch Billing - Lose Resets? — r/OpenAI
  3. What would make you actually use a personal AI assistant everyday? — r/artificial
  4. My ComfyUI Nodes and Workflows - Krea 2 (Turbo and Raw), Z-Image (Turbo and Base), MiniMax Music 3, Image2Text and LLM Chat (with Tools), Torch, Apple MLX and Cloud — r/comfyui
  5. MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash — r/LocalLLaMA
  6. Running un-filtered / NSFW models or workflows in ComfyUI on an Apple Silicon Mac (16GB RAM) - Recommendations? — r/comfyui
  7. SynthID available globally starting today — r/StableDiffusion
  8. Apple Vision vs MobileSAM in a Mac annotation tool: 67 ms to the first mask in a small test — r/computervision

Get the daily brief of stories like this at 6:30 every morning →