AINewsnow

Deepseek training 2T and plans 8T model

Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model. https://x.com/wallstengine/status/2101982843656388644 Current DeepSeek models: Flash parameter count of 552 billion Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per to…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-21 12:38 · r/LocalLLaMA
    Deepseek training 2T and plans 8T model

More stories

  1. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI
  2. ChatGPT Plus feels overkill for my usage — any good coding agents under $20? — r/AI_Agents
  3. Foulmouth Qwen 3.8 27b, an unexpected thought process... — r/LocalLLaMA
  4. Is there any use of a local llm with a 20B LLM? — r/AI_Agents
  5. Ai used for chatbots — r/artificial
  6. I gave 6 different AIs the same 5 questions — r/AI_Agents
  7. PromptDeck v1.1.0 – open-source desktop app to benchmark local AND cloud LLMs side-by-side (Ollama, LM Studio + OpenRouter, Groq, DeepSeek…) — r/LocalLLaMA
  8. What’s your favorite AI model for coding right now? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →