AINewsnow

Gated Segmented State Space — attention replacement that beats a param-matched Transformer on quality, speed AND memory (full code)

One night, six experiments (V1–V6), one Colab T4. I ripped self-attention out of a decoder-only Transformer and replaced it with a gated linear recurrence over a fixed 256-dim state: Dynamic selective gate: g_t = σ(W_g x_t + b_g) — per-token/channel learned filter Hard reset mask: state zeroed at n…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-10-08 00:44 · r/learnmachinelearning
    Gated Segmented State Space — attention replacement that beats a param-matched Transformer on quality, speed AND memory (full code)

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. Introducing Mistral Large 4 — Mistral AI News
  4. Sharing AI progress in mathematics — OpenAI News
  5. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  6. Everything announced at Microsoft's Windows and Surface event — Engadget
  7. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  8. Introducing Playground: Create and play custom games — Google AI Blog

Get the daily brief of stories like this at 6:30 every morning →