Qwen3.8-Flash-Next 125B running locally on a Strix Halo mini-PC: 43 tok/s, tool calls 2× faster, KL divergence 0.116 vs full precision. The model diagnosed and wrote one of the runtime fixes itself.
Coverage of "Qwen3.8-Flash-Next 125B running locally on a Strix Halo mini-PC: 43 tok/s, tool calls 2× faster, KL divergence 0.116 vs full precision. The model diagnosed and wrote one of the runtime fixes itself." from 1 source, with a live timeline of who reported what and when.
Read the full story at r/LocalLLM ↗