基于深度强化学习算法的多无人水面航行器编队构造

展开
  • (大连海事大学 航海学院,辽宁 大连 116026) 
关巍*(1982 — ),男,博士,教授,博士生导师。研究方向:船舶智能航行决策,船舶运动控制。张诚(1996 — ),男,硕士生,研究方向:船舶编队控制。崔哲闻(1999 — ),男,硕士生,研究方向:船舶智能航行决策。韩虎生(1997 — ),男,硕士生,研究方向:船舶智能航行决策。 E-mail:zhangcheng1108@dlmu.edu.cn。

网络出版日期: 2024-02-20

基金资助

国家自然科学基金(52171342)

Formation Construction of Multiple Unmanned Surface Vehicles Based on Deep Reinforcement Learning Algorithm

Expand
  • (Navigation College, Dalian Maritime University, Dalian 116026, China)

Online published: 2024-02-20

摘要

伴随科技的快速发展,多无人船系统在军事、营救以及护送等任务场景中展现出巨大的潜力。研究旨在基于多智能体深度强化学习算法,对多无人水面航行器系统的编队构造问题进行探究。针对传统多智能体深度确定性策略梯度算法(MADDPG)收敛速度较慢的问题,研究通过在值函数阶段引入了注意力机制来提升多无人水面航行器系统的编队决策模型的收敛速度,并通过无人水面航行器的编队模型与编队避碰和编队构造奖励函数的配合,提升了多无人水面航行器完成编队构造任务的效率。最后,通过仿真结果验证了本文所提出的方法可有效地完成多种环境下的多无人水面航行器编队构造任务,为未来多无人船编队构造应用提供理论研究基础。

本文引用格式

关巍, 张诚, 崔哲闻, 韩虎生 . 基于深度强化学习算法的多无人水面航行器编队构造[J]. 大连海事大学学报, 2025 , 51(1) : 11 -20 . DOI: 10.16411/j.cnki.issn1006-7736.2025.01.002

Abstract

With the rapid development of science and technology, multi-unmanned ship systems have shown great potential in military, rescue and escort mission scenarios. The purpose of this paper is to explore the formation construction problem of multiple unmanned surface vehicle systems based on multi-agent deep reinforcement learning algorithm. Considering the sluggish convergence speed of the conventional multi-agent deep deterministic policy gradient algorithm (MADDPG), this study incorporates the attention mechanism into the value function stage to enhance the convergence speed of the formation decision model for a multi-UAV system. Through the cooperation of the formation model of the unmanned surface vehicle with the formation collision avoidance and the formation construction reward function, the efficiency of the multi-UAV to complete the formation construction task is finally improved. The simulation results conclusively demonstrate the efficacy of the proposed method in accomplishing multi-unmanned surface vehicle formation construction tasks, thereby establishing a solid theoretical foundation for future applications of multi-unmanned ship formation construction.

参考文献

[1]訾润.无人机编队在多点火情救援中的任务分配研究[J].中国新技术新产品, 2022(012):000.
Zi Run. Research on Task Allocation of Drone Formation in Multi point Fire Rescue [J]. China NewTechnology and Products, 2022 (012): 000. (in Chinese)
[2]X Li, Y Lin, Z Du, R Yin and C Wu, "Multi-robot Cooperative Transport Simulation System," 2023 IEEE/CIC International Conference on Communications in China (ICCC Workshops), Dalian, China, 2023:1-6.
[3]Wang D, Fu M. Adaptive Formation Control for Waterjet USV With Input and Output Constraints Based on Bioinspired Neurodynamics[J].IEEE Access, 2019, PP(99):1-1. 
[4]Dai S L, He S, Luo F, et al. Leader-follower formation control of fully actuated USVs with prescribed performance and collision avoidance[C]//2017 36th Chinese Control Conference (CCC). 2017.
[5]郑延斌,席鹏雪,王林林,等.基于人工势场法的多智能体编队避障方法[J].计算机应用, 2018, 38(12):6.
Zheng Yanbin, Xi Pengxue, Wang Linlin, et al. Obstacle avoidance method for multi-agent formation based on artificial potential field method [J]. Computer Application, 2018, 38 (12): 6. (in Chinese)
[6]童进军,何黎明,田作华. 船舶动力定位系统的数学模型[J]. 船舶工程, 2002(5): 27-29. 
Tong Jinjun, He Liming, Tian Zuohua Mathematical Model of Ship Dynamic Positioning System [J]. Ship Engineering, 2002 (5): 27-29. (in Chinese)
[7]王醒策,张汝波,顾国昌. 多机器人动态编队的强化学习算法研究[J]. 计算机研究与发展, 2003, 40(10):7.
Wang Xingce, Zhang Rubo, Gu Guochang Research on Reinforcement Learning Algorithm for Dynamic Formation of Multiple Robots [J] Computer Research and Development, 2003, 40 (10): 7. (in Chinese)
[8]Xie L , Wang S , Markham A ,et al.Towards Monocular Vision based Obstacle Avoidance through Deep Reinforcement Learning.2017[2023-09-07].
[9]Zhao Y, Ma Y, Hu S. USV Formation and Path-Following Control via Deep Reinforcement LearningWith Random Braking[J]. IEEE Transactions on Neural Networks and Learning Systems, 2021, PP(99): 1-11. 
[10]Lin Y, Wang M, Zhou X, et al. Dynamic Spectrum Interaction of UAV Flight Formation Communication With Priority: A Deep Reinforcement Learning Approach[J]. IEEE Transactions on Cognitive Communications and Networking, 2020, PP(99):1-1.
[11]Sui Z, Pu Z, Yi J, et al. Formation Control with Collision Avoidance through Deep Reinforcement Learning[C]//2019 International Joint Conference on Neural Networks (IJCNN). 2019. 
[12]Zhao J, Member, Li F, et al. Deep Reinforcement Learning based Model-free On-line Dynamic Multi-Microgrid Formation to Enhance Resilience[J]. 2022
[13]范培潇,柯松,杨军,等.基于改进多智能体深度确定性策略梯度的多微网负荷频率协同控制策略[J].电网技术, 2022(009): 046.
Fan Peixiao, Ke Song, Yang Jun, et al. Multi microgrid load frequency collaborative control strategy based on improved multi-agent deep deterministic strategy gradient [J]. Power Grid Technology, 2022 (009): 046. (in Chinese)
[14]W Guan, Z Cui, and X Zhang. Intelligent Smart Marine Autonomous Surface Ship Decision System Based on Improved PPO Algorithm,” Sensors,2022,22(15):5732.
[15]K Miyazaki, N Matsunaga, and K Murata, "Formation path learning for cooperative transportation ofmultiple robots using MADDPG," 2021 21st International Conference on Control, Automation and Systems (ICCAS), Jeju, Korea, Republic of, 2021, pp.1619-1623.

文章导航

/