Skip to main navigation Skip to search Skip to main content

Implicit Differentiation for Fast Hyperparameter Selection in Non-Smooth Convex Learning

  • Université Paris-Saclay
  • Centre de Recherches de Climatologie, CNRS UMR 5210, Université de Bourgogne
  • University of Genoa
  • Brain team
  • University of Montpellier (UMR MiVEGEC)

Research output: Contribution to journalArticlepeer-review

20 Citations (Scopus)

Abstract

Finding the optimal hyperparameters of a model can be cast as a bilevel optimization problem, typically solved using zero-order techniques. In this work we study first-order methods when the inner optimization problem is convex but non-smooth. We show that the forward-mode differentiation of proximal gradient descent and proximal coordinate descent yield sequences of Jacobians converging toward the exact Jacobian. Using implicit differentiation, we show it is possible to leverage the non-smoothness of the inner problem to speed up the computation. Finally, we provide a bound on the error made on the hypergradient when the inner optimization problem is solved approximately. Results on regression and classification problems reveal computational benefits for hyperparameter optimization, especially when multiple hyperparameters are required.

Original languageEnglish
JournalJournal of Machine Learning Research
Volume23
Publication statusPublished - 1 Apr 2022
Externally publishedYes

Keywords

  • Convex optimization
  • Lasso
  • bilevel optimization
  • generalized linear models
  • hyperparameter optimization
  • hyperparameter selection

Fingerprint

Dive into the research topics of 'Implicit Differentiation for Fast Hyperparameter Selection in Non-Smooth Convex Learning'. Together they form a unique fingerprint.

Cite this