TY - GEN
T1 - A Lagrangian Framework for Safe Cooperative Reinforcement Learning
AU - Das, Soham
AU - Chamon, Luiz F.O.
AU - Paternain, Santiago
AU - Eksin, Ceyhun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025/1/1
Y1 - 2025/1/1
N2 - We consider the problem of safe cooperative multiagent reinforcement learning (MARL) within the framework of a constrained multiagent Markov decision process (MDP). Agents share a common value function and learn to coordinate their actions to maximize a joint objective while adhering to system-level constraints. These constraints can enforce safety, reliability, or additional regulatory requirements governing the evolution of the multiagent system. We propose a Lagrangian-based approach, where agents iteratively solve a relaxed Lagrangian MDP using a joint learning mechanism. During execution, agents independently follow their policies, accumulating constraint violations over an epoch, which are then used to update the Lagrange multipliers. We show that continuous execution of this primal-dual algorithm produces episodes which are feasible almost surely. Further, we prove that the sequence of policies generated by the algorithm yields a nonstationary approximately optimal solution for the safe cooperative MARL problem.
AB - We consider the problem of safe cooperative multiagent reinforcement learning (MARL) within the framework of a constrained multiagent Markov decision process (MDP). Agents share a common value function and learn to coordinate their actions to maximize a joint objective while adhering to system-level constraints. These constraints can enforce safety, reliability, or additional regulatory requirements governing the evolution of the multiagent system. We propose a Lagrangian-based approach, where agents iteratively solve a relaxed Lagrangian MDP using a joint learning mechanism. During execution, agents independently follow their policies, accumulating constraint violations over an epoch, which are then used to update the Lagrange multipliers. We show that continuous execution of this primal-dual algorithm produces episodes which are feasible almost surely. Further, we prove that the sequence of policies generated by the algorithm yields a nonstationary approximately optimal solution for the safe cooperative MARL problem.
UR - https://www.scopus.com/pages/publications/105031875087
U2 - 10.1109/CDC57313.2025.11312681
DO - 10.1109/CDC57313.2025.11312681
M3 - Conference contribution
AN - SCOPUS:105031875087
T3 - Proceedings of the IEEE Conference on Decision and Control
SP - 5112
EP - 5119
BT - 2025 IEEE 64th Conference on Decision and Control, CDC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 64th IEEE Conference on Decision and Control, CDC 2025
Y2 - 9 December 2025 through 12 December 2025
ER -