TY - GEN
T1 - Deep Reinforcement Learning Approach for UAV Search Path Planning In Discrete Time and Space
AU - Benalaya, Najoua
AU - Amdouni, Ichrak
AU - Adjih, Cedric
AU - Laouiti, Anis
AU - Saidane, Leila Azouz
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024/1/1
Y1 - 2024/1/1
N2 - Path planning for search missions carried out by Unmanned Aerial Vehicles (UAVs) is a challenging problem. This is due to UAV limited energy budget and the importance of time for search operations. The objective of this study is to come up with an approach to minimize the total search time required to locate a specific target. To achieve this, we deployed a deep reinforcement learning (DRL) model based on the Proximal Policy Optimization (PPO) algorithm to solve the combinatorial optimization problem of UAV search path planning within a minimized search time. A smart reward formulation is designed to achieve the learning goal, fulfill the search requirement, and encourage the agent to select search paths that minimize search time. In addition, we employed Optuna hyperparameter optimization framework to systematically select optimal parameters for the PPO model. Most importantly, thanks to the state representation we considered, the model is generalized and adaptable to various search environments. The PPO model succeeds to compute an accurate search path to be followed by the UAV searcher. Results of the model are compared with results previously obtained with a linear program. We found that the PPO achieves almost the same expected search time, which proves the great relevance of the reward design and the hyperparameters selection we made.
AB - Path planning for search missions carried out by Unmanned Aerial Vehicles (UAVs) is a challenging problem. This is due to UAV limited energy budget and the importance of time for search operations. The objective of this study is to come up with an approach to minimize the total search time required to locate a specific target. To achieve this, we deployed a deep reinforcement learning (DRL) model based on the Proximal Policy Optimization (PPO) algorithm to solve the combinatorial optimization problem of UAV search path planning within a minimized search time. A smart reward formulation is designed to achieve the learning goal, fulfill the search requirement, and encourage the agent to select search paths that minimize search time. In addition, we employed Optuna hyperparameter optimization framework to systematically select optimal parameters for the PPO model. Most importantly, thanks to the state representation we considered, the model is generalized and adaptable to various search environments. The PPO model succeeds to compute an accurate search path to be followed by the UAV searcher. Results of the model are compared with results previously obtained with a linear program. We found that the PPO achieves almost the same expected search time, which proves the great relevance of the reward design and the hyperparameters selection we made.
KW - Deep Reinforcement Learning
KW - Hyperparameters Search
KW - Optuna
KW - PPO
KW - Reward Design
KW - Search Path Planning
KW - UAVs
U2 - 10.1109/IWCMC61514.2024.10592510
DO - 10.1109/IWCMC61514.2024.10592510
M3 - Conference contribution
AN - SCOPUS:85199985871
T3 - 20th International Wireless Communications and Mobile Computing Conference, IWCMC 2024
SP - 1437
EP - 1442
BT - 20th International Wireless Communications and Mobile Computing Conference, IWCMC 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 20th IEEE International Wireless Communications and Mobile Computing Conference, IWCMC 2024
Y2 - 27 May 2024 through 31 May 2024
ER -