Formation Construction of Multiple Unmanned Surface Vehicles Based on Deep Reinforcement Learning Algorithm

Expand
  • (Navigation College, Dalian Maritime University, Dalian 116026, China)

Online published: 2024-02-20

Abstract

With the rapid development of science and technology, multi-unmanned ship systems have shown great potential in military, rescue and escort mission scenarios. The purpose of this paper is to explore the formation construction problem of multiple unmanned surface vehicle systems based on multi-agent deep reinforcement learning algorithm. Considering the sluggish convergence speed of the conventional multi-agent deep deterministic policy gradient algorithm (MADDPG), this study incorporates the attention mechanism into the value function stage to enhance the convergence speed of the formation decision model for a multi-UAV system. Through the cooperation of the formation model of the unmanned surface vehicle with the formation collision avoidance and the formation construction reward function, the efficiency of the multi-UAV to complete the formation construction task is finally improved. The simulation results conclusively demonstrate the efficacy of the proposed method in accomplishing multi-unmanned surface vehicle formation construction tasks, thereby establishing a solid theoretical foundation for future applications of multi-unmanned ship formation construction.

Cite this article

GUAN Wei, ZHANG Cheng, CUI Zhe-wen, Han Hu-sheng . Formation Construction of Multiple Unmanned Surface Vehicles Based on Deep Reinforcement Learning Algorithm[J]. Journal of Dalian Maritime University, 2025 , 51(1) : 11 -20 . DOI: 10.16411/j.cnki.issn1006-7736.2025.01.002

References

[1]訾润.无人机编队在多点火情救援中的任务分配研究[J].中国新技术新产品, 2022(012):000.
Zi Run. Research on Task Allocation of Drone Formation in Multi point Fire Rescue [J]. China NewTechnology and Products, 2022 (012): 000. (in Chinese)
[2]X Li, Y Lin, Z Du, R Yin and C Wu, "Multi-robot Cooperative Transport Simulation System," 2023 IEEE/CIC International Conference on Communications in China (ICCC Workshops), Dalian, China, 2023:1-6.
[3]Wang D, Fu M. Adaptive Formation Control for Waterjet USV With Input and Output Constraints Based on Bioinspired Neurodynamics[J].IEEE Access, 2019, PP(99):1-1. 
[4]Dai S L, He S, Luo F, et al. Leader-follower formation control of fully actuated USVs with prescribed performance and collision avoidance[C]//2017 36th Chinese Control Conference (CCC). 2017.
[5]郑延斌,席鹏雪,王林林,等.基于人工势场法的多智能体编队避障方法[J].计算机应用, 2018, 38(12):6.
Zheng Yanbin, Xi Pengxue, Wang Linlin, et al. Obstacle avoidance method for multi-agent formation based on artificial potential field method [J]. Computer Application, 2018, 38 (12): 6. (in Chinese)
[6]童进军,何黎明,田作华. 船舶动力定位系统的数学模型[J]. 船舶工程, 2002(5): 27-29. 
Tong Jinjun, He Liming, Tian Zuohua Mathematical Model of Ship Dynamic Positioning System [J]. Ship Engineering, 2002 (5): 27-29. (in Chinese)
[7]王醒策,张汝波,顾国昌. 多机器人动态编队的强化学习算法研究[J]. 计算机研究与发展, 2003, 40(10):7.
Wang Xingce, Zhang Rubo, Gu Guochang Research on Reinforcement Learning Algorithm for Dynamic Formation of Multiple Robots [J] Computer Research and Development, 2003, 40 (10): 7. (in Chinese)
[8]Xie L , Wang S , Markham A ,et al.Towards Monocular Vision based Obstacle Avoidance through Deep Reinforcement Learning.2017[2023-09-07].
[9]Zhao Y, Ma Y, Hu S. USV Formation and Path-Following Control via Deep Reinforcement LearningWith Random Braking[J]. IEEE Transactions on Neural Networks and Learning Systems, 2021, PP(99): 1-11. 
[10]Lin Y, Wang M, Zhou X, et al. Dynamic Spectrum Interaction of UAV Flight Formation Communication With Priority: A Deep Reinforcement Learning Approach[J]. IEEE Transactions on Cognitive Communications and Networking, 2020, PP(99):1-1.
[11]Sui Z, Pu Z, Yi J, et al. Formation Control with Collision Avoidance through Deep Reinforcement Learning[C]//2019 International Joint Conference on Neural Networks (IJCNN). 2019. 
[12]Zhao J, Member, Li F, et al. Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience[J]. 2022
[13]范培潇,柯松,杨军,等.基于改进多智能体深度确定性策略梯度的多微网负荷频率协同控制策略[J].电网技术, 2022(009): 046.
Fan Peixiao, Ke Song, Yang Jun, et al. Multi microgrid load frequency collaborative control strategy based on improved multi-agent deep deterministic strategy gradient [J]. Power Grid Technology, 2022 (009): 046. (in Chinese)
[14]W Guan, Z Cui, and X Zhang. Intelligent Smart Marine Autonomous Surface Ship Decision System Based on Improved PPO Algorithm,” Sensors,2022,22(15):5732.
[15]K Miyazaki, N Matsunaga, and K Murata, "Formation path learning for cooperative transportation ofmultiple robots using MADDPG," 2021 21st International Conference on Control, Automation and Systems (ICCAS), Jeju, Korea, Republic of, 2021, pp.1619-1623.

Outlines

/