Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training platform. The goal is to keep healthy workers training when another pipeline stage goes offline. Crucible combines data-parallel replicas with pipeline parallelism. Each replica holds a copy of the mo…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-22 15:47 · r/MachineLearning
Simulating fault tolerance with stage skipping in pipeline-parallel training [R]