基于Dueling-DDQN的船舶智能避碰决策方法

展开
  • (大连海事大学 航海学院,辽宁 大连 116026)
关巍*(1982 — ),男,教授,博士生导师,研究方向:船舶运动控制理论,船舶智能航行避障决策,E-mail:gwwtxdy@dlmu.edu.cn

网络出版日期: 2024-12-27

基金资助

国家自然科学基金资助项目(52171342)

Intelligent collision avoidance decision-making method for ships based on Dueling-DDQN

Expand
  • (Navigation College, Dalian Maritime University, Dalian 116026, China)

Online published: 2024-12-27

摘要

针对由全球海上船舶数量增长导致的碰撞事故频发问题,提出一种基于深度强化学习(DRL)算法的船舶智能避碰决策模型,基于对抗双深度Q学习(Dueling-DDQN)与船舶领域模型的建立,设计奖励函数时充分考虑了《国际海上避碰规则》(COLREGs)及船舶偏航等要素,以确保避碰决策的合规性与合理性。搭建仿真环境模拟多船会遇场景,并利用神经网络模型处理复杂环境信息,进行模型训练与验证。实验结果显示,相比传统深度Q学习算法,本文模型在收敛速度和稳定性方面均表现出显著优势,能够准确判断会遇局面,并依据COLREGs采取恰当的避碰措施,展现出较高的决策准确性和可靠性,可为船舶在复杂海况下的智能航行提供有效的决策支持。

本文引用格式

关巍, 王淼淼, 韩虎生, 崔哲闻, 蔡珊珊 . 基于Dueling-DDQN的船舶智能避碰决策方法[J]. 大连海事大学学报, 2024 , 50(4) : 22 -30 . DOI: 10.16411/j.cnki.issn1006-7736.2024.04.003

Abstract

A ship intelligent collision avoidance decision model based on deep reinforcement learning (DRL) algorithm was proposed to address the frequent collision accidents caused by the global increase in the number of ships at sea. The model was based on adversarial dual deep Q-learning (Dueling DDQN) and the establishment of ship domain models, and when designing the reward function, factors such as COLREGs (International Regulations for Preventing Collisions at Sea) and ship deviation were fully considered to ensure the compliance and rationality of collision avoidance decisions. A simulation environment was constructed to simulate the scenario of multiple ships encountering, and neural network models were used to process complex environmental information for model training and validation. Experimental results show that compared with traditional deep Q-learning algorithms, the model proposed in this paper exhibits significant advantages in convergence speed and stability, which can accurately determine the encounter situation and take appropriate collision avoidance measures based on COLREGs, demonstrating high decision accuracy and reliability. It can provide effective decision support for intelligent navigation of ships in complex sea conditions.

参考文献

印桂生,孙颖,杨东梅,等. 船舶安全事故典型案例分析与研究[J]. 船舶, 2021, 32(4): 1-14.
YIN G S, SUN Y, YANG D M, et al. Analysis and research on typical cases of ship safety accidents[J]. Ship & Boat, 2021, 32(4): 1-14. (in Chinese)
[2]李永杰,张瑞,魏慕恒,等. 船舶自主航行关键技术研究现状与展望[J]. 中国舰船研究, 2021, 16(1): 32-44.
LI Y J, ZHANG R, WEI M H, et al. Current status and prospects of key technologies for autonomous navigation of ships[J]. Chinese Journal of Ship Research, 2021, 16(1): 32-44. (in Chinese)
[3]HE Z B, LIU C G, CHU X M, et al. Dynamic anti-collision A-star algorithm for multi-ship encounter situations[J]. Applied Ocean Research, 2022, 118: 102995.
[4]刘渐道,刘文,张英俊,等. 基于改进动态窗口法的无人水面艇自主避碰算法[J]. 上海海事大学学报, 2021, 42(2): 1-7.
LIU J D, LIU W, ZHANG Y J, et al. An autonomous collision avoidance algorithm for unmanned surface vehicles based on an improved dynamic window approach[J]. Journal of Shanghai Maritime University, 2021, 42(2): 1-7. (in Chinese)
[5]GUAN W, WANG K. Autonomous collision avoidance of unmanned surface vehicles based on improved A-star and dynamic window approach algorithms[J]. IEEE Intelligent Transportation Systems Magazine, 2023, 15(3): 36-50.
[6]ZHANG X T, CHEN X Y. Path planning method for unmanned surface vehicle based on RRT* and DWA[C]// International Conference on Multimedia Technology and Enhanced Learning, ICMTEL 2021.[S.l.:s.n.], 2021: 518-527.
[7]HE Z B, CHU X M, LIU C G, et al. A novel model predictive artificial potential field based ship motion planning method considering COLREGs for complex encounter scenarios[J]. ISA Transactions, 2023, 134: 58-73.
[8]PEROLAT J, DE VYLDER B, HENNES D, et al. Mastering the game of stratego with model-free multiagent reinforcement learning[J]. Science, 2022, 378(6623): 990-996.
[9]SILVER D, HUBERT T, SCHRITTWIESER J, et al. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play[J]. Science, 2018, 362(6419): 1140-1144.
[10]SHEN H Q, HASHIMOTO H, MATSUDA A, et al. Automatic collision avoidance of multiple ships based on deep Q-learning[J]. Applied Ocean Research, 2019, 86: 268-288.
[11]XU X L, LU Y, LIU G, et al. COLREGs-abiding hybrid collision avoidance algorithm based on deep reinforcement learning for USVs[J]. Ocean Engineering, 2022, 247: 110749.
[12]ZHAO Y M, HAN F L, HAN D F, et al. Decision-making for the autonomous navigation of USVs based on deep reinforcement learning under IALA maritime buoyage system[J]. Ocean Engineering, 2022, 266: 112557.
[13]CUI Z W, GUAN W, ZHANG X K, et al. Autonomous navigation decision-making method for a smart marine surface vessel based on an improved soft Actor-Critic algorithm[J]. Journal of Marine Science and Engineering, 2023, 11(8): 1554.
[14]WU C B, YU W N, LI G Z, et al. Deep reinforcement learning with dynamic window approach based collision avoidance path planning for maritime autonomous surface ships[J]. Ocean Engineering, 2023, 284: 115208.
[15]关巍,崔哲闻,罗文哲. 基于改进PPO算法的船舶自主避碰决策方法[J]. 大连海事大学学报, 2023, 49(4): 28-36.
GUAN W, CUI Z W, LUO W Z. A ship autonomous collision avoidance decision-making method based on improved PPO algorithm[J]. Journal of Dalian Maritime University, 2023, 49(4): 28-36. (in Chinese)
[16]关巍,罗文哲,崔哲闻. 基于深度强化学习的无人驾驶船舶避碰行为决策[J]. 大连海事大学学报, 2024, 50(1): 11-19.
GUAN W, LUO W Z, CUI Z W.Collision avoidance behavior decision-making of unmanned ship based on deep reinforcement learning[J]. Journal of Dalian Maritime University, 2024, 50(1): 11-19. (in Chinese)
[17]FOSSEN T I. Handbook of marine craft hydrodynamics and motion control[M]. [S.l.]: John Wiley and Sons, 2011.
[18]WANG N. An intelligent spatial collision risk based on the quaternion ship domain[J]. Journal of Navigation, 2010, 63(4): 733-749.
[19]VAN HASSELT H, GUEZ A, SILVER D. Deep reinforcement learning with double Q-Learning[C]//30th AAAI Conference on Artificial Intelligence, AAAI 2016. [S.l.:s.n.],2016: 2094-2100.
[20]MNIH V, KAVUKCUOGLU K, SILVER D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540): 529-533.

文章导航

/