Regularization in Deep Nets
A modern network routinely has more parameters than training examples — by classical statistical intuition, this should guarantee catastrophic overfitting. It usually doesn't, and the regularisation techniques on this page are less about shrinking weights (as in Overfitting and Regularization's classical L1/L2 story) and more about injecting noise or stopping early.