Whallm 1.1.11: Swift1.5-Qwen3.8 support, and Qwen3.8 now runs at 15–17 tok/s decode and 500–600 tok/s prefill at 16K on a 64 GB Mac
Whallm is a free, open-source macOS app that runs large MoE models on Apple Silicon by streaming experts from your SSD. New in 1.1.11 Swift1.5-Qwen3.8-Flash-Next support , with text and image input. Decode: ~9 → 15–17 tok/s. MTP is now on by default. Prefill: ~146 → 500–600 tok/s at 16K input. Peak…
Read the full story at r/LocalLLM ↗