AINewsnow

MTP + prefix caching still broken on Arc Pro B70: llm-scaler 0.26.0-b2's release notes claim a fix, my tests disagree — filed as intel/llm-scaler#741

**Why I'm posting this:** I run Qwen3.8-27B on 2x Arc Pro B70 as the coder tier of a home-lab agent fleet. For agentic workloads, prefix caching matters more than speculative decoding — every turn re-sends a big shared prefix, so cache hits are the real win and MTP's decode gains are secondary. Whe…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-03 14:46 · r/LocalLLM
    MTP + prefix caching still broken on Arc Pro B70: llm-scaler 0.26.0-b2's release notes claim a fix, my tests disagree — filed as intel/llm-scaler#741

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  4. Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
  5. Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
  6. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  7. The latest AI news we announced in September 2026 — Google Gemini Blog
  8. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)

Get the daily brief of stories like this at 6:30 every morning →