Skip to main navigation Skip to search Skip to main content

Model-free Posterior Sampling via Learning Rate Randomization

  • Daniil Tiapkin
  • , Denis Belomestny
  • , Daniele Calandriello
  • , Éric Moulines
  • , Remi Munos
  • , Alexey Naumov
  • , Pierre Perrault
  • , Michal Valko
  • , Pierre Ménard
  • Ecole Polytechnique
  • National Research University
  • University of Duisburg-Essen
  • Google DeepMind
  • Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
  • IDEMIA
  • Ecole Normale Supérieure de Lyon

Research output: Contribution to journalConference articlepeer-review

Abstract

In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the best of our knowledge, RandQL is the first tractable model-free posterior sampling-based algorithm. We analyze the performance of RandQL in both tabular and non-tabular metric space settings. In tabular MDPs, RandQL achieves a regret bound of order Oe(√H5SAT), where H is the planning horizon, S is the number of states, A is the number of actions, and T is the number of episodes. For a metric state-action space, RandQL enjoys a regret bound of order Oe(H5/2T(dz+1)/(dz+2)), where dz denotes the zooming dimension. Notably, RandQL achieves optimistic exploration without using bonuses, relying instead on a novel idea of learning rate randomization. Our empirical study shows that RandQL outperforms existing approaches on baseline exploration environments.

Original languageEnglish
JournalAdvances in Neural Information Processing Systems
Volume36
Publication statusPublished - 1 Jan 2023
Externally publishedYes
Event37th Conference on Neural Information Processing Systems, NeurIPS 2023 - New Orleans, United States
Duration: 10 Dec 202316 Dec 2023

Fingerprint

Dive into the research topics of 'Model-free Posterior Sampling via Learning Rate Randomization'. Together they form a unique fingerprint.

Cite this