基于改进PPO算法的船舶自主避碰决策

展开
  •  (大连海事大学 航海学院,辽宁 大连 116026)
关巍*(1982 — ),男,博士,教授,博士生导师,研究方向:船舶运动控制理论,船舶自主避障决策,小型智能航行系统。崔哲闻(1999 — ),男,硕士生,研究方向:船舶自主避障决策。罗文哲(1998 — ),男,硕士生,研究方向:船舶自主避障决策。E-mail:gwwtxdy@dlmu.edu.cn.
关巍*(1982 — ),男,博士,教授,博士生导师,研究方向:船舶运动控制理论,船舶自主避障决策,小型智能航行系统。E-mail:gwwtxdy@dlmu.edu.cn

收稿日期: 2023-05-30

  修回日期: 2023-07-18

  录用日期: 2023-07-18

  网络出版日期: 2023-07-18

基金资助

国家自然科学基金资助项目(52171342)

Ship autonomous collision avoidance decision based on improved PPO algorithm

Expand
  • (Navigation College, Dalian Maritime University, Dalian 116026, China)

Received date: 2023-05-30

  Revised date: 2023-07-18

  Accepted date: 2023-07-18

  Online published: 2023-07-18

摘要

为减少船舶避碰决策过程中因人为失误导致的海难事故,提出一种基于改进近端策略优化(PPO)算法的船舶自主避碰决策。在传统PPO算法广义优势估计基础上加入自适应基线调整,并且使用长短期记忆网络(LSTM)改进网络结构。船舶的航行信息和激光雷达矢量线被应用于神经网络的输入,航行制导、角度偏差以及《1972年避碰规则》均被纳入改进的奖励函数设计。两船和多船会遇场景仿真实验表明:本文提出的避碰决策可使船舶实现自主航行,并在避碰过程中符合《避碰规则》,为处理复杂局面下的船舶避碰决策提供了参考。

本文引用格式

关巍, 崔哲闻, 罗文哲 . 基于改进PPO算法的船舶自主避碰决策[J]. 大连海事大学学报, 2023 , 49(4) : 28 -36 . DOI: 10.16411/j.cnki.issn1006-7736.2023.04.004

Abstract

To prevent shipwreck accidents caused by human errors in the process of collision avoidance decision-making, a ship autonomous collision avoidance decision-making method based on improved proximal policy optimization (PPO) algorithm was proposed. Based on the generalized advantage estimation of the traditional PPO algorithm, the adaptive baseline adjustment was added and long short-term memory network was used to improve the network structure. In addition, the ship navigation information and the lidar vector line were used as the input of the neural network. Navigation guidance, angle deviation and the International Regulations for Preventing Collisions at Sea (COLREGs) were designed into the improved reward function. The simulation experiments of two or multiple ship encounter scenarios show that the collision avoidance decision-making method proposed can enable ships to achieve autonomous navigation and comply with the COLREGs during the collision avoidance process, while provides a reference for handling ship collision avoidance in complex situations.

参考文献

[1]Guan W, Peng H W, Zhang X K, et al. Ship Steering Adaptive CGS Control Based on EKF Identification Method[J]. Journal of Marine Science and Engineering, 2022, 10(2).
[2]Cheng X, Liu Z Y, Soc I C. Trajectory optimization for ship navigation safety using genetic annealing algorithm[C]//3rd International Conference on Natural Computation (ICNC 2007), 2007: 385.
[3]Phanthong T, Maki T, Ura T, et al. Application of A* algorithm for real-time path re-planning of an unmanned surface vehicle avoiding underwater obstacles[J]. Journal of Marine Science of Application, 2014, 13(1): 12.
[4]Liu X Y, Li Y, Zhang J, et al. Self-Adaptive Dynamic Obstacle Avoidance and Path Planning for USV Under Complex Maritime Environment[J]. IEEE Access, 2019, 7: 114945-114954.
[5]吕红光,尹 勇,尹建川. 混合智能系统在船舶自动避碰决策中的应用[J]. 大连海事大学学报, 2015, 41(4): 29-36.
LV Hong-guang,YIN Yong,YIN Jian-chuan. Application of hybrid intelligent systems in automatic collision avoidance decision-making for ships[J]. Journal of Dalian Maritime University, 2015, 41(4): 29-36. (in Chinese)
[6]苗鹏,刘克中,辛旭日,陈逸涵,吴晓烈. 基于改进NSGA-Ⅱ的船舶避碰决策辅助算法[J]. 大连海事大学学报, 2021, 47(4): 10-18.
MIAO Peng,LIU Ke-zhong,XIN Xu-ri,CHEN Yi-han,WU Xiao-lie. Ship collision avoidance decision aids based on improved NSGA-Ⅱ[J]. Journal of Dalian Maritime University, 2021, 47(4): 10-18. (in Chinese)
[7]宁君, 黄寓旸, 尤恽, 李伟, 高鉴一, 张帅. 基于混合粒子群算法的船舶避碰决策[J]. 大连海事大学学报, 2023, 49(1): 34-43.
NING Jun, HUANG Yuyang, YOU Yun, LI Wei, GAO Jianyi, ZHANG Shuai. Ship collision avoidance decision based on hybrid particle swarm algorithm[J]. Journal of Dalian Maritime University, 2023, 49(1): 34-43. (in Chinese)
[8]杨柏丞,赵志垒. 基于改进模拟退火算法的多船会遇避碰决策[J].大连海事大学学报,2018,44(2):22-26.
YANG B C, ZHAO Z L. Multi-ship encounter collision avoidance decisions based on improved simulated annealing algorithm [J]. Journal of Dalian Maritime University, 2018,44(2):22-26. (in Chinese)
[9]Shen H, Hashimoto H, Matsuda A, et al. Automatic collision avoidance of multiple ships based on deep Q-learning[J]. Applied Ocean Research, 2019, 86: 268-288.
[10]Xu X L, Lu Y, Liu X C, et al. Intelligent collision avoidance algorithms for USVs via deep reinforcement learning under COLREGs[J]. Ocean Engineering, 2020, 217.
[11]Sawada R, Sato K, Majima T. Automatic ship collision avoidance using deep reinforcement learning with LSTM in continuous action spaces[J]. Journal of Marine Science and Technology, 2021, 26(2): 509-524.
[12]Guo S Y, Zhang X G, Zheng Y S, et al. An Autonomous Path Planning Model for Unmanned Ships Based on Deep Reinforcement Learning[J]. Sensors,2020, 20(2):426.
[13]T. I. Fossen, Handbook of Marine Craft Hydrodynamics and Motion Control: Handbook of Marine Craft Hydrodynamics and Motion Control, 2011. 
[14]Śmierzchalski R. Ships' domains as collision risk at sea in the evolutionary method of trajectory planning[J], In: Saeed, K., Pejaś, J. (eds) Information Processing and Security Systems. Springer, Boston, MA, 2015.
[15]Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O., Proximal Policy Optimization Algorithms. [J] Computer science, 2017.
[16]Guan W, Cui Z, Zhang X. Intelligent Smart Marine Autonomous Surface Ship Decision System Based on Improved PPO Algorithm[J]. Sensors, 2022, 22(15): 5732.
[17]Cui Z, Guan W, Luo W. Intelligent Ship Decision System Based on DDPG Algorithm[J]. 2022 International Conference on Computer Engineering and Artificial Intelligence (ICCEAI), 2022: 700-705.
[18]B. Bingham, C. Aguero, M. Mccarrin et al., "Toward Maritime Robotic Simulation in Gazebo." MTS/IEEE OCEANS Conference, Oct, 2019. 

文章导航

/