AINewsnow

Deepseek V4.1 Flash is 748B, not 552B

This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.

People keep on getting confused about this, so I looked at the safetensors on hf. The title should have been "Deepseek V4.1 Flash is 748B total/552B base, not 284B or 305B or 485B or 522B" The model is not 284B. The original Deepseek V4 Flash is 284B, but not the V4.1 Flash model The model is not 3…

Read the full story at r/LocalLLaMA ↗

Timeline · 16 reports

  1. 2026-09-13 06:12 · r/reinforcementlearning
    DeepSeek-V4.1-Flash Tech Report
  2. 2026-09-12 19:09 · r/huggingface
    We’re testing DeepSeek V4.1 flash bs GLM 5.3 flash. V4.1 is free to use.
  3. 2026-09-12 17:12 · r/LocalLLM
    DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
  4. 2026-09-12 05:56 · Latent Space
    [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
  5. 2026-09-11 22:24 · Unite.AI
    Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context
  6. 2026-09-11 14:56 · r/huggingface
    DeepSeek-V4.1-Flash: GGUF + 4.75bpw EXL3 are out, looking for devs with 4× DGX Sparks to help validate the EXL3 TP4 recipe
  7. 2026-09-11 03:09 · Pandaily
    Cambricon Day-0 Adapts DeepSeek-V4.1-Flash on vLLM Stack
  8. 2026-09-10 20:00 · r/LocalLLaMA
    CPU Only Experimental Sloppy Deepseek V4.1 Flash
  9. 2026-09-10 19:45 · r/LocalLLaMA
    Livebench added Deepseek v4.1 flash
  10. 2026-09-10 19:05 · r/LocalLLM
    DeepSeek V4.1 Flash is 510 GB but only about 150 of it has to be in memory. I read the shard headers and made a fit checker.
  11. 2026-09-10 16:05 · r/huggingface
    DeepSeek V4.1 Flash is available in HuggingChat
  12. 2026-09-10 16:05 · r/LocalLLaMA
    DeepSeek V4.1 Flash is available in HuggingChat
  13. 2026-09-10 11:23 · r/LocalLLaMA
    Deepseek V4.1 Flash Release Video [Made with Deepseek V4.1 Flash]
  14. 2026-09-10 10:01 · r/LocalLLaMA
    guide to using reasoning_effort on deepseek v4.1 flash
  15. 2026-09-10 08:37 · r/LocalLLaMA
    DeepSeek-V4.1-Flash surprised ....
  16. 2026-09-10 08:27 · r/LocalLLaMA
    Deepseek V4.1 Flash is 748B, not 552B

More stories

  1. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  2. Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash — r/LocalLLaMA
  3. DeepSeek’s Insane New Architecture — Two Minute Papers
  4. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  5. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  6. 2 TB of Cheep Pmem200 Dimms can Run Kimi K3 at tg128 ~ 1 t/s · pp512 5.6558 — r/LocalLLM
  7. Are we over-engineering AI agent workflows? — r/AI_Agents
  8. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI

Get the daily brief of stories like this at 6:30 every morning →