Loss-spike gating: a measured negative result
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
Loss spikes are a known hazard in LLM pretraining. PaLM 540B hit about 20 of them and handled each by rolling back ~100 steps and skipping 200-500 data batches. The obvious response is a detector: watch the gradient, catch the spike, skip the update before it corrupts the optimizer state. I tried t…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-06 09:10 · r/deeplearning
Loss-spike gating: a measured negative result