How to prevent overfitting with a custom optimizer?
I've designed a custom optimizer that outperforms ADAM on small models. In XOR, it reduces the loss to 0 in 3 steps. In MNISt, it reduces the loss to 1e-6 in 200 steps, while having higher accuracy than ADAM in 2000 steps. But when I try it in nanogpt, even though it's 5x faster than ADAM, it insta…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-28 21:13 · r/learnmachinelearning
How to prevent overfitting with a custom optimizer?