Passer à la navigation principale Passer à la recherche Passer au contenu principal

Reinforcement Learning with a Terminator

  • Guy Tennenholtz
  • , Nadav Merlis
  • , Lior Shani
  • , Shie Mannor
  • , Uri Shalit
  • , Gal Chechik
  • , Assaf Hallak
  • , Gal Dalal
  • Technion - Israel Institute of Technology
  • NVIDIA
  • Bar-Ilan University

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

2 Citations (Scopus)

Résumé

We present the problem of reinforcement learning with exogenous termination. We define the Termination Markov Decision Process (TerMDP), an extension of the MDP framework, in which episodes may be interrupted by an external non-Markovian observer. This formulation accounts for numerous real-world situations, such as a human interrupting an autonomous driving agent for reasons of discomfort. We learn the parameters of the TerMDP and leverage the structure of the estimation problem to provide state-wise confidence bounds. We use these to construct a provably-efficient algorithm, which accounts for termination, and bound its regret. Motivated by our theoretical analysis, we design and implement a scalable approach, which combines optimism (w.r.t. termination) and a dynamic discount factor, incorporating the termination probability. We deploy our method on high-dimensional driving and MinAtar benchmarks. Additionally, we test our approach on human data in a driving setting. Our results demonstrate fast convergence and significant improvement over various baseline approaches.

langue originaleAnglais
titreAdvances in Neural Information Processing Systems 35 - 36th Conference on Neural Information Processing Systems, NeurIPS 2022
rédacteurs en chefS. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, A. Oh
EditeurNeural information processing systems foundation
ISBN (Electronique)9781713871088
étatPublié - 1 janv. 2022
Modification externeOui
Evénement36th Conference on Neural Information Processing Systems, NeurIPS 2022 - New Orleans, États-Unis
Durée: 28 nov. 20229 déc. 2022

Série de publications

NomAdvances in Neural Information Processing Systems
Volume35
ISSN (imprimé)1049-5258

Une conférence

Une conférence36th Conference on Neural Information Processing Systems, NeurIPS 2022
Pays/TerritoireÉtats-Unis
La villeNew Orleans
période28/11/229/12/22

Empreinte digitale

Examiner les sujets de recherche de « Reinforcement Learning with a Terminator ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation