Passer à la navigation principale Passer à la recherche Passer au contenu principal

Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training

  • Meta Ai
  • École Polytechnique
  • Université Paris Dauphine

Résultats de recherche: Contribution à un journalArticle de conférenceRevue par des pairs

1 Citation (Scopus)

Résumé

Training Deep Neural Networks (DNNs) with small batches using Stochastic Gradient Descent (SGD) often results in superior test performance compared to larger batches. This implicit bias is attributed to the specific noise structure inherent to SGD. When ensuring Differential Privacy (DP) in DNNs’ training, DP-SGD adds Gaussian noise to the clipped gradients. However, large-batch training still leads to a significant performance decrease, posing a challenge as strong DP guarantees necessitate the use of massive batches. Our study first demonstrates that this phenomenon extends to Noisy-SGD (DP-SGD without clipping), suggesting that the stochasticity, not the clipping, is responsible for this implicit bias, even with additional isotropic Gaussian noise. We then theoretically analyze the solutions obtained with continuous versions of Noisy-SGD for the Linear Least Square and Diagonal Linear Network settings. Our analysis reveals that the additional noise indeed amplifies the implicit bias. It suggests that the performance issues of private training stem from the same underlying principles as SGD, offering hope for improvements in large batch training strategies.

langue originaleAnglais
Pages (de - à)3295-3303
Nombre de pages9
journalProceedings of Machine Learning Research
Volume238
étatPublié - 1 janv. 2024
Evénement27th International Conference on Artificial Intelligence and Statistics, AISTATS 2024 - Valencia, Espagne
Durée: 2 mai 20244 mai 2024

Empreinte digitale

Examiner les sujets de recherche de « Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation