大连海事大学学报 >
2021 , Vol. 47 >Issue 2: 1 - 10
DOI: https://doi.org/10.16411/j.cnki.issn1006-7736.2021.02.001
基于强化学习的指定性能轨迹跟踪最优控制
收稿日期: 2020-12-03
修回日期: 2021-04-04
网络出版日期: 2021-04-04
基金资助
国家自然科学基金资助项目(51579024);辽宁省“兴辽英才计划”资助项目(XLYC1807013);辽宁省高等学校创新人才支持计划(LR2017024);水下机器人技术国防科技重点实验室稳定支持课题资助项目(SXJQR2018WDKT03);中央高校基本科研业务费专项基金资助项目(3132019318; 3132019344).
Optimal control of specified performance trajectory tracking based on reinforcement learning
Received date: 2020-12-03
Revised date: 2021-04-04
Online published: 2021-04-04
为保证无人船对期望轨迹的高动态精确跟踪,针对带有输入饱和限制的无人船轨迹跟踪系统,提出一种基于强化学习的指定性能轨迹跟踪最优控制策略.采用双曲正切函数逼近饱和函数,并利用指定性能控制技术将无人船跟踪误差进行误差转换,使跟踪动态误差约束在期望的指定范Χ内以保证轨迹跟踪系统的高动态性能.在此基础上,采用最优控制理论与强化学习中基于神经网络的ActorCritic控制框架设计无人船轨迹跟踪最优控制方法,同时解决了控制输入饱和限制的问题.基于李雅普诺夫稳定性理论,证明了本文控制策略下无人船轨迹跟踪系统的稳定性,保证了跟踪误差均在指定的范Χ内.仿真实验也验证了本文策略的有效性与优越性.
杨忱 , 赵红 , 王宁 , 高颖 , 郭晨 . 基于强化学习的指定性能轨迹跟踪最优控制[J]. 大连海事大学学报, 2021 , 47(2) : 1 -10 . DOI: 10.16411/j.cnki.issn1006-7736.2021.02.001
In order to ensure the high dynamic and accurate tracking of the desired trajectory, an optimal trajectory tracking control strategy with specified performance based on reinforcement learning was proposed for the an unmanned surface vehicle (USV) trajectory tracking system with input saturation limit. The hyperbolic tangent function was used to approximate the saturation function, and by the specified performance control technique, the tracking error of the unmanned ship was transformed, so that the tracking dynamic error was constrained within the expected specified range to ensure the high dynamic performance of the trajectory tracking system. On this basis, the optimal control theory and the actor critical control framework based on neural network in reinforcement learning were used to design the optimal control method of unmanned ship trajectory tracking, and the problem of control input saturation limit was solved. Based on the Lyapunov stability theory, the stability of the proposed trajectory tracking system under the control strategy was proved, and the tracking error was guaranteed within the specified range. Simulation results also verify the effectiveness and superiority of the proposed strategy.
/
| 〈 |
|
〉 |