Will Nvidia Vera Rubin actually make LLM pre-training faster? And are 10T+ parameter models next?
Nvidia's Vera Rubin numbers look huge, but most of the headline gains seem tied to low-precision formats and inference. How much of that actually carries over to pre-training? Also curious whether we'll start seeing 10T+ parameter models, or if data, power, and cost are the real limits now, and the…
Read the full story at r/ArtificialInteligence ↗
Timeline · 1 report
- 2026-09-29 14:34 · r/ArtificialInteligence
Will Nvidia Vera Rubin actually make LLM pre-training faster? And are 10T+ parameter models next?