AINewsnow

DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI went after, and they cut theirs to 890 bytes per token. They released DeepSeek-V4.1-Flash, a 552B MoE model with 1M-token context that activates only 8B parameters per token d…

Read the full story at r/machinelearningnews ↗

Timeline · 1 report

  1. 2026-09-10 07:54 · r/machinelearningnews
    DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

More stories

  1. Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash — r/LocalLLaMA
  2. DeepSeek’s Insane New Architecture — Two Minute Papers
  3. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  4. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  5. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  6. I enjoyed the daily HF papers today — r/LocalLLaMA
  7. Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading — SemiAnalysis
  8. Coming soon...... Optimized for DEEPSEEK Flash.... Though model Agnostic.... message me to test.... cem888.ai — r/huggingface

Get the daily brief of stories like this at 6:30 every morning →