Passer à la navigation principale Passer à la recherche Passer au contenu principal

Reinforcement Learning with Trajectory Feedback

  • Yonathan Efroni
  • , Nadav Merlis
  • , Shie Mannor
  • Technion - Israel Institute of Technology
  • Microsoft Research
  • NVIDIA

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

The standard feedback model of reinforcement learning requires revealing the reward of every visited state-action pair. However, in practice, it is often the case that such frequent feedback is not available. In this work, we take a first step towards relaxing this assumption and require a weaker form of feedback, which we refer to as trajectory feedback. Instead of observing the reward obtained after every action, we assume we only receive a score that represents the quality of the whole trajectory observed by the agent, namely, the sum of all rewards obtained over this trajectory. We extend reinforcement learning algorithms to this setting, based on least-squares estimation of the unknown reward, for both the known and unknown transition model cases, and study the performance of these algorithms by analyzing their regret. For cases where the transition model is unknown, we offer a hybrid optimistic-Thompson Sampling approach that results in a tractable algorithm.

langue originaleAnglais
titre35th AAAI Conference on Artificial Intelligence, AAAI 2021
EditeurAssociation for the Advancement of Artificial Intelligence
Pages7288-7295
Nombre de pages8
ISBN (Electronique)9781713835974
Les DOIs
étatPublié - 1 janv. 2021
Modification externeOui
Evénement35th AAAI Conference on Artificial Intelligence, AAAI 2021 - Virtual, Online
Durée: 2 févr. 20219 févr. 2021

Série de publications

Nom35th AAAI Conference on Artificial Intelligence, AAAI 2021
Volume8B

Une conférence

Une conférence35th AAAI Conference on Artificial Intelligence, AAAI 2021
La villeVirtual, Online
période2/02/219/02/21

Empreinte digitale

Examiner les sujets de recherche de « Reinforcement Learning with Trajectory Feedback ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation