AINewsnow

focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737)

I forked llama.cpp to implement Declarative Attention (arXiv:2609.02737, Google DeepMind and KAIST AI). The model declares in its own output which context chunks it needs ), and the engine listens and restricts what the following tokens can attend to. No scorer, no training: just prompting plus an…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-20 14:07 · r/LocalLLaMA
    focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737)

More stories

  1. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  2. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  3. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  4. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  5. Google’s Gemini AI hacked into other companies, adding to ‘rogue’ AI incidents — Washington Post AI
  6. We need to talk about Irregular — r/singularity
  7. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  8. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →