a 56 layer network that fits its own training data worse than a 20 layer one, same recipe same seed
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
i was writing the ResNet chapter of a pytorch book and i did not want to just tell the reader that a deep plain network gets worse, i wanted to actually watch it happen, so i trained four networks on CIFAR-10 with one recipe and one seed for all of them, and the only things i changed are the depth…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-08-20 10:44 · r/deeplearning
a 56 layer network that fits its own training data worse than a 20 layer one, same recipe same seed