Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Aug 25, 2020

Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, Jascha Sohl-Dickstein

Figure 1 for Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Figure 2 for Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Figure 3 for Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Figure 4 for Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Share this with someone who'll enjoy it:

Abstract:Machine learning is predicated on the concept of generalization: a model achieving low error on a sufficiently large training set should also perform well on novel samples from the same distribution. We show that both data whitening and second order optimization can harm or entirely prevent generalization. In general, model training harnesses information contained in the sample-sample second moment matrix of a dataset. For a general class of models, namely models with a fully connected first layer, we prove that the information contained in this matrix is the only information which can be used to generalize. Models trained using whitened data, or with certain second order optimization schemes, have less access to this information; in the high dimensional regime they have no access at all, producing models that generalize poorly or not at all. We experimentally verify these predictions for several architectures, and further demonstrate that generalization continues to be harmed even when theoretical requirements are relaxed. However, we also show experimentally that regularized second order optimization can provide a practical tradeoff, where training is still accelerated but less information is lost, and generalization can in some circumstances even improve.

* 15+7 pages, 7 figures; added references, edited model descriptions for clarity, results unchanged

View paper on

OpenReview

Share this with someone who'll enjoy it:

Title:Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible

Paper and Code