Passer à la navigation principale Passer à la recherche Passer au contenu principal

Learning Safe Policies via Primal-Dual Methods

  • School of Engineering and Applied Science

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

23 Citations (Scopus)

Résumé

In this paper, we study the learning of safe policies in the setting of reinforcement learning problems. This is, we aim to control a Markov Decision Process (MDP) of which we do not know the transition probabilities, but we have access to sample trajectories through experiments. We define safety as the agent remaining in a desired safe set with high probability for every time instance. We therefore consider a constrained MDP where the constraints are probabilistic. Due to the difficulty of addressing these constraints in a reinforcement learning framework, we propose an ergodic relaxation of the problem. Nonetheless, this relaxation is such that we are able to provide safety guarantees on the resulting policies. To compute these policies, we resource to a stochastic primal-dual method. We test the proposed approach in a navigation task in a grid world. The numerical results show that our algorithm is capable of dynamically adapting the policy to the environment and the required safety levels.

langue originaleAnglais
titre2019 IEEE 58th Conference on Decision and Control, CDC 2019
EditeurInstitute of Electrical and Electronics Engineers Inc.
Pages6491-6497
Nombre de pages7
ISBN (Electronique)9781728113982
Les DOIs
étatPublié - 1 déc. 2019
Modification externeOui
Evénement58th IEEE Conference on Decision and Control, CDC 2019 - Nice, France
Durée: 11 déc. 201913 déc. 2019

Série de publications

NomProceedings of the IEEE Conference on Decision and Control
Volume2019-December
ISSN (imprimé)0743-1546
ISSN (Electronique)2576-2370

Une conférence

Une conférence58th IEEE Conference on Decision and Control, CDC 2019
Pays/TerritoireFrance
La villeNice
période11/12/1913/12/19

Empreinte digitale

Examiner les sujets de recherche de « Learning Safe Policies via Primal-Dual Methods ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation