Résumé
In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the best of our knowledge, RandQL is the first tractable model-free posterior sampling-based algorithm. We analyze the performance of RandQL in both tabular and non-tabular metric space settings. In tabular MDPs, RandQL achieves a regret bound of order Oe(√H5SAT), where H is the planning horizon, S is the number of states, A is the number of actions, and T is the number of episodes. For a metric state-action space, RandQL enjoys a regret bound of order Oe(H5/2T(dz+1)/(dz+2)), where dz denotes the zooming dimension. Notably, RandQL achieves optimistic exploration without using bonuses, relying instead on a novel idea of learning rate randomization. Our empirical study shows that RandQL outperforms existing approaches on baseline exploration environments.
| langue originale | Anglais |
|---|---|
| journal | Advances in Neural Information Processing Systems |
| Volume | 36 |
| état | Publié - 1 janv. 2023 |
| Modification externe | Oui |
| Evénement | 37th Conference on Neural Information Processing Systems, NeurIPS 2023 - New Orleans, États-Unis Durée: 10 déc. 2023 → 16 déc. 2023 |
Empreinte digitale
Examiner les sujets de recherche de « Model-free Posterior Sampling via Learning Rate Randomization ». Ensemble, ils forment une empreinte digitale unique.Contient cette citation
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver