AINewsnow

Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K

I ran with sparse attention on for the better part of a month before I sat down and benchmarked it at short context, and it had been costing me time that whole stretch without me noticing. There are two separate things in the engine that both get called sparse attention, and only one of them is the…

Read the full story at r/machinelearningnews ↗

Timeline · 1 report

  1. 2026-10-02 07:24 · r/machinelearningnews
    Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K

More stories

  1. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  4. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  5. Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
  6. OpenAI Delays Release of Latest Model Over Safety Concerns — Wired AI
  7. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  8. Introducing GPT-6.1 Sol — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →