大连海事大学学报 >
2022 , Vol. 48 >Issue 3: 20 - 30
DOI: https://doi.org/10. 16411 / j. cnki. issn1006-7736. 2022. 03. 003
基于强化学习的自学习遗传算法在船舶调度中的应用
收稿日期: 2022-04-13
修回日期: 2022-04-29
网络出版日期: 2022-04-29
基金资助
国家自然科学基金面上项目( 51779028);中央引导地方科技发展资金自由探索类基础研究项目(2021Szvup014)
Application of self-learning genetic algorithm based on reinforcement learning in ship scheduling
Received date: 2022-04-13
Revised date: 2022-04-29
Online published: 2022-04-29
为有效调度进出港船舶以解决港口拥堵问题,提出一种基于强化学习的自学习遗传算法( GA-RL),在 GA- RL 中,以遗传算法为基本优化模型,利用 Q-learning 算法自适应调整交叉和变异参数来提高算法的搜索能力;同时,构建可动态调参的马尔科夫决策过程(MDP)模型,在MDP 模型中,为全面评估种群性能,提出基于种群适应度函数的状态集,并设计了有效减少目标值的奖励机制;最后,以黄骅港综合港区为例,选取不同组算例进行仿真实验,结果验证了模型和算法的有效性,该方法可显著减少船舶在港等待时间并提升港口通航效率。
关键词: 船舶调度; 自学习遗传算法; Q-learning; 单向航道
李润佛 , 张新宇 , 李俊杰 , 姜玲玲 . 基于强化学习的自学习遗传算法在船舶调度中的应用[J]. 大连海事大学学报, 2022 , 48(3) : 20 -30 . DOI: 10. 16411 / j. cnki. issn1006-7736. 2022. 03. 003
s to solve the p roblem of port congestion , a self-learning genetic algorithm based on reinforce ment learning (GA- RL ) was proposed. In GA-RL, taking the genetic algorithm as the b asi c optimizat ion m odel , and Q-learn ing algo-rithm was use d to ada ptive ly adju st the c ros sove r and mutat ion par ameter s to improve the search ability of the algo rithm. At the same time, a dynamically adjustable Markov decision pr ocess (M DP) model was construct ed. In the MD P model, in ord er to c omp rehens ivel y e valuate the popula tion perf orm- ance, a st ate se t based on the pop ulat ion fitness function was propo sed, a nd a reward me chanism for effe ctiv ely reducing the targ et va lue was desig ned. Fi nal ly, taking the comp rehens ive port area of H uanghua p ort as an exam ple, different group s of exam ples w ere selected for s imu latio n exper imen ts. The re- sults verif y the effect iveness of the m odel an d alg orit hm, whi ch can si gnif icantly red uc e the waiting time of ships in port and improve the efficiency of port navigatio n.
/
| 〈 |
|
〉 |