MTP + prefix caching still broken on Arc Pro B70: llm-scaler 0.26.0-b2's release notes claim a fix, my tests disagree — filed as intel/llm-scaler#741
**Why I'm posting this:** I run Qwen3.8-27B on 2x Arc Pro B70 as the coder tier of a home-lab agent fleet. For agentic workloads, prefix caching matters more than speculative decoding — every turn re-sends a big shared prefix, so cache hits are the real win and MTP's decode gains are secondary. Whe…
Read the full story at r/LocalLLM ↗