Deepseek training 2T and plans 8T model
Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model. https://x.com/wallstengine/status/2101982843656388644 Current DeepSeek models: Flash parameter count of 552 billion Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per to…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-21 12:38 · r/LocalLLaMA
Deepseek training 2T and plans 8T model