AINewsnow

Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

DeepSeek-V4.1-Flash is available now on Baseten Model APIs, Baseten announced on September 11, 2026, bringing the 552B-parameter multimodal mixture-of-experts (MoE) model, which pairs 8B active parameters for prefill with 16B for decode across a 1M-token context window, to the inference provider's…

Read the full story at Unite.AI ↗

Timeline · 13 reports

  1. 2026-09-14 16:00 · AlphaSignal
    What DeepSeek-V4.1-Flash teaches us about efficient AI
  2. 2026-09-14 12:00 · KDnuggets
    Why DeepSeek-V4.1-Flash Is Such an Exciting Open Model Release
  3. 2026-09-14 03:12 · r/LocalLLM
    Whallm now supports DeepSeek V4.1 Flash on Apple Silicon — plus performance updates and built-in benchmarks
  4. 2026-09-14 02:04 · r/learnmachinelearning
    Why DeepSeek V4.1 Flash reconstructs part of its KV cache from only 128 tokens
  5. 2026-09-14 01:37 · r/LocalLLaMA
    DeepSeek V4.1 Flash beats Astra on AA's new benchmark
  6. 2026-09-13 16:03 · r/LocalLLaMA
    Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes
  7. 2026-09-13 13:55 · AlphaSignal
    DeepSeek V4.1-Flash Stripped of Safety Hits 100% Harmful Prompt Compliance
  8. 2026-09-13 11:02 · TheSequence
    The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough
  9. 2026-09-13 06:12 · r/reinforcementlearning
    DeepSeek-V4.1-Flash Tech Report
  10. 2026-09-12 19:09 · r/huggingface
    We’re testing DeepSeek V4.1 flash bs GLM 5.3 flash. V4.1 is free to use.
  11. 2026-09-12 17:12 · r/LocalLLM
    DeepSeek V4.1 Flash running locally on 8× A40 — ~40 tok/s Q2_K, ~32 tok/s Q4_K_M
  12. 2026-09-12 05:56 · Latent Space
    [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
  13. 2026-09-11 22:24 · Unite.AI
    Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context

More stories

  1. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  2. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  3. Google’s Gemini AI hacked into other companies, adding to ‘rogue’ AI incidents — Washington Post AI
  4. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  5. Week in review: OpenAI ships managed Agents API, Apple's new Siri reportedly runs on Gemini, and three vendors add agent spend controls — r/artificial
  6. Gemini 4 หลุด แต่ paper ที่ Google เพิ่งตีพิมพ์ตรวจสอบได้ทุกตัวเลข — DEV Community — AI
  7. Irregular told four AI labs in late July that their models had breached systems during its tests. The public learned in stages, and Google went last. — The Next Web
  8. Local LLM on iPhone 18 is impressive — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →